هل يمكن لنماذج اللغة الصغيرة والكبيرة القادرة على التفكير تقييم مقالات المجلات من حيث جودة البحث، وهل تساعد المتوسطات والدروس القليلة؟
Can small and reasoning large language models score journal articles for research quality and do averaging and few-shot help?

شارك:
المجلة: Scientometrics
DOI: https://doi.org/10.1007/s11192-026-05585-2
تاريخ النشر: 2026-03-02
المؤلف: Mike Thelwall وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم

نظرة عامة

تقيّم هذه الدراسة قدرة نماذج اللغة الكبيرة (LLMs) المتوسطة والصغيرة على تقييم جودة المقالات العلمية في مجالات الطب والصحة وعلوم الحياة. أشارت الدراسات السابقة إلى أن نماذج مثل ChatGPT وGemini تتوافق بشكل معتدل مع تقييمات الخبراء. تختبر الدراسة الحالية نماذج إضافية، بما في ذلك متغيرات Gemma3 وLlama4 Scout وQwen3 وMagistral Small وDeepSeek R1، باستخدام مجموعة بيانات تضم 2,780 مقالة عبر ستة مجالات. تكشف النتائج أن نماذج LLM المتوسطة تؤدي بشكل مشابه للنماذج الأكبر، بينما يمكن للنماذج الأصغر التي تحتوي على أكثر من 4 مليارات معلمة أيضًا تقييم جودة البحث بفعالية، خاصة عند استخدام تقنيات متوسط الدرجات. لم تُظهر نماذج التفكير ميزة كبيرة، وتقترح الدراسة أن التحفيز القليل قد يقدم بعض الفوائد، على الرغم من أن تنوع التحفيز قد يكون أكثر أهمية.

تؤكد النتائج على جدوى استخدام نماذج LLM لتقييم البحث، حيث تظهر نماذج متعددة قدرات مشابهة، مما يجعلها مناسبة للاستخدام في بيئات تعاني من قيود الموارد. ومع ذلك، تشير الدراسة إلى أن درجات LLM لا تتماشى تمامًا مع تقييمات البشر، وغالبًا ما تتجمع حول قيم أعلى وتظهر عدم موثوقية في تقييم المقالات الفردية. وبالتالي، يُوصى باستخدام نماذج LLM بشكل أساسي لترتيب المقالات بدلاً من تقديم درجات جودة نهائية. تؤكد الدراسة على الحاجة إلى الحذر في تفسير مخرجات LLM وتقترح أن هذه الدرجات يجب أن تكمل حكم الخبراء بدلاً من استبداله، خاصة بالنظر إلى تعقيدات توافق الخبراء في تقييمات الجودة.

مقدمة

تناقش مقدمة هذه الورقة البحثية إمكانيات نماذج اللغة الكبيرة (LLMs) كبدائل لمؤشرات الاقتباس التقليدية لتقييم جودة البحث وتأثيره. تبرز النتائج التي تشير إلى وجود ارتباطات كبيرة بين مخرجات نماذج LLM البارزة، مثل ChatGPT-4o وGemini 1.5 Flash، وتقييمات الخبراء لجودة البحث، حيث تتفوق الأولى على مقاييس الاقتباس في العديد من المجالات. كما تشير الورقة إلى ارتباطات معتدلة لنموذج LLM ذو الأوزان المفتوحة Gemma3 27b، بينما أظهرت نماذج الأوزان المفتوحة الأخرى، مثل Qwen2.5-72b وLlama3، ارتباطات ضعيفة فقط في سياق بيولوجي طبي محدد.

يحدد المؤلفون أربعة أسئلة بحثية رئيسية تهدف إلى استكشاف أداء نماذج LLM المتوسطة في تقييم البحث الأكاديمي، والعلاقة بين حجم LLM وقدرته على التقييم، وفعالية استراتيجيات التحفيز القليل، وكفاءة متوسط الدرجات عبر تقييمات متعددة. تعترف الدراسة بالقيود في نطاقها، مستبعدة النماذج القديمة وبعض التباينات في المعلمات، مع التركيز على إمكانيات نماذج التفكير وتأثيرها على تقييمات جودة البحث.

طرق البحث

في هذه الدراسة، طور المؤلفون مجموعة بيانات شاملة تضم مقالات من مجالات متنوعة، تم تقييم كل منها من قبل خبراء لتوليد درجات مقابلة. كان الهدف الرئيسي هو تقييم أداء نماذج اللغة الكبيرة المختلفة (LLMs) واستراتيجيات التحفيز في تقييم هذه المقالات، مع التركيز على ربط الدرجات المعينة من قبل الخبراء البشر بتلك التي تولدها نماذج LLM.

مع الاعتراف بحدود نماذج LLM، خاصة ميلها لإنتاج درجات متجمعة، أكدت الدراسة على ارتباط الرتب كمقياس حاسم. تتيح هذه الطريقة فهمًا أكثر دقة لكيفية قرب ترتيب المقالات وفقًا لتقييمات الخبراء من الترتيبات المستمدة من مخرجات LLM، مما يوفر رؤى حول موثوقية وفعالية الذكاء الاصطناعي في سياقات التقييم الأكاديمي.

النتائج

يقدم قسم “النتائج” في الورقة البحثية النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يحدد النتائج الأساسية، بما في ذلك البيانات الإحصائية، والاتجاهات الملحوظة، وأي ارتباطات هامة تم تحديدها. غالبًا ما يتم توضيح النتائج من خلال الجداول أو الرسوم البيانية أو الأشكال، مما يوفر تمثيلًا بصريًا ويسهل فهم البيانات.

يؤكد القسم على تداعيات النتائج، مناقشًا كيف تساهم في المعرفة الحالية في هذا المجال. يتم الإبلاغ عن مقاييس محددة، مثل قيم p أو فترات الثقة، لدعم صحة النتائج. بشكل عام، يخدم هذا القسم لتسليط الضوء على أهمية النتائج بالنسبة للأسئلة البحثية المطروحة والفرضيات المختبرة.

المناقشة

تسلط قسم المناقشة في الورقة الضوء على القيود والتداعيات لنتائج الدراسة المتعلقة بتقييم جودة البحث في مجالات الصحة وعلوم الحياة باستخدام نماذج اللغة الكبيرة (LLMs). تبني الدراسة على دراسات سابقة استخدمت إما مجموعات بيانات صغيرة أو اعتمدت على درجات متوسطة قسمية كبدائل لدرجات جودة المقالات الفردية. تركز الدراسة الحالية على عينة من 500 مقالة من ست وحدات تقييم (UoAs) ضمن إطار التميز البحثي في المملكة المتحدة (REF) 2021، مستخدمة كل من درجات التجميع القسمية ودرجات جودة المقالات الفردية كمعايير ذهبية للتثليث. تشير النتائج إلى أنه بينما يمكن أن يقلل استخدام درجات المتوسطات القسمية من قوة الارتباط، فإنه يسمح بإجراء تحليل مقارن بين نماذج LLM المختلفة، على الرغم من وجود تحيزات محتملة بسبب تأثيرات القسم.

تشير النتائج إلى أن نماذج LLM المتوسطة تؤدي بشكل مشابه للنماذج الأكبر المستندة إلى السحابة، دون وجود تفوق واضح. تظهر استراتيجية التحفيز القليل وعدًا، مما يعزز الارتباط في عدة حالات، على الرغم من أن بعض النماذج تواجه صعوبة مع هذه الطريقة. من المهم أن متوسط الدرجات من عدة نماذج LLM يحقق باستمرار ارتباطات أعلى من الاعتماد على درجات فردية. ومع ذلك، تعترف الدراسة بعدة قيود، بما في ذلك مجموعة البيانات التي تركز على المملكة المتحدة، والطبيعة الذاتية لدرجات الجودة الفردية، وإمكانية أن تؤثر الاختلافات المعمارية بين نماذج LLM على الأداء. بشكل عام، بينما توفر الدراسة رؤى قيمة حول تطبيق نماذج LLM لتقييم البحث، فإنها تؤكد أيضًا على الحاجة إلى الحذر في تفسير النتائج بسبب التعقيدات والقيود الكامنة في المنهجيات المستخدمة.

Journal: Scientometrics
DOI: https://doi.org/10.1007/s11192-026-05585-2
Publication Date: 2026-03-02
Author(s): Mike Thelwall et al.
Primary Topic: Artificial Intelligence in Healthcare and Education

Overview

This research evaluates the ability of various medium-sized and smaller Large Language Models (LLMs) to assess the quality of journal articles in medical, health, and life sciences. Previous studies indicated that LLMs like ChatGPT and Gemini correlate moderately with expert ratings. The current study tests additional models, including Gemma3 variants, Llama4 Scout, Qwen3, Magistral Small, and DeepSeek R1, using a dataset of 2,780 articles across six fields. The findings reveal that medium-sized LLMs perform comparably to larger models, while smaller models with over 4 billion parameters can also effectively rate research quality, particularly when employing score averaging techniques. Reasoning models did not demonstrate a significant advantage, and the study suggests that few-shot prompting may offer some benefit, although the variety of prompts could be more crucial.

The results underscore the practicality of using LLMs for research evaluation, as multiple models exhibit similar capabilities, making them suitable for deployment in resource-constrained environments. However, the study notes that LLM scores do not align perfectly with human evaluations, often clustering around higher values and demonstrating unreliability for individual article assessments. Consequently, LLMs are recommended primarily for ranking articles rather than providing definitive quality scores. The research emphasizes the need for caution in interpreting LLM outputs and suggests that these scores should complement expert judgment rather than replace it, particularly given the complexities of expert agreement in quality assessments.

Introduction

The introduction of this research paper discusses the potential of Large Language Models (LLMs) as alternatives to traditional citation-based indicators for assessing research quality and impact. It highlights findings that indicate substantial correlations between the outputs of prominent LLMs, such as ChatGPT-4o and Gemini 1.5 Flash, and expert judgments on research quality, with the former outperforming citation metrics in many fields. The paper also notes moderate correlations for the open weights LLM Gemma3 27b, while other open weights models, like Qwen2.5-72b and Llama3, demonstrated only weak correlations in a specific biomedical context.

The authors outline four primary research questions aimed at exploring the performance of medium-sized LLMs in academic research scoring, the relationship between LLM size and scoring ability, the effectiveness of few-shot prompting strategies, and the efficacy of averaging scores across multiple evaluations. The study acknowledges limitations in its scope, excluding older models and certain parameter variations, while focusing on the potential of reasoning models and their impact on research quality evaluations.

Methods

In this study, the authors developed a comprehensive dataset comprising articles from various disciplines, each evaluated by experts to generate corresponding scores. The primary objective was to assess the performance of different large language models (LLMs) and prompting strategies in scoring these articles, with a focus on correlating the scores assigned by human experts with those generated by the LLMs.

Recognizing the limitations of LLMs, particularly their tendency to produce clustered scores, the research emphasized rank correlation as a critical metric. This approach allows for a more nuanced understanding of how closely the ordering of articles by expert evaluations aligns with the rankings derived from LLM outputs, thereby providing insights into the reliability and effectiveness of AI in academic scoring contexts.

Results

The “Results” section of the research paper presents key findings derived from the conducted experiments or analyses. It outlines the primary outcomes, including statistical data, observed trends, and any significant correlations identified. The results are often illustrated through tables, graphs, or figures, which provide visual representation and facilitate comprehension of the data.

The section emphasizes the implications of the findings, discussing how they contribute to the existing body of knowledge in the field. Specific metrics, such as p-values or confidence intervals, are reported to substantiate the validity of the results. Overall, this section serves to highlight the importance of the findings in relation to the research questions posed and the hypotheses tested.

Discussion

The discussion section of the paper highlights the limitations and implications of the study’s findings regarding the evaluation of research quality in the health and life sciences using large language models (LLMs). The research builds on previous studies that have either utilized small datasets or relied on departmental mean scores as proxies for individual article quality scores. The current study focuses on a sample of 500 articles from six Units of Assessment (UoAs) within the UK’s Research Excellence Framework (REF) 2021, employing both departmental aggregate scores and individual article-level quality scores as gold standards for triangulation. The findings suggest that while the use of departmental mean scores can dampen correlation strength, it allows for comparative analysis between different LLMs, albeit with potential biases due to departmental influences.

The results indicate that medium-sized LLMs perform comparably to larger cloud-based models, with no clear superiority established. The few-shot prompting strategy shows promise, enhancing correlation in several instances, although some models struggle with this approach. Importantly, averaging scores from multiple LLMs consistently yields higher correlations than relying on single scores. However, the study acknowledges several limitations, including the UK-centric dataset, the subjective nature of the individual quality scores, and the potential for architectural differences among LLMs to affect performance. Overall, while the study provides valuable insights into the application of LLMs for research assessment, it also underscores the need for caution in interpreting the results due to the inherent complexities and limitations of the methodologies employed.

شارك: