هل OpenAlex مناسب لتقييم جودة البحث وأي مؤشر اقتباس هو الأفضل؟
Is OpenAlex suitable for research quality evaluation and which citation indicator is best?

شارك:
المجلة: Journal of the Association for Information Science and Technology، المجلد: 76، العدد: 12
DOI: https://doi.org/10.1002/asi.70020
تاريخ النشر: 2025-08-22
المؤلف: Mike Thelwall وآخرون
الموضوع الرئيسي: البحث في علم القياسات العلمية والبيبلومetrics

نظرة عامة

هذا القسم من ورقة البحث يقيم فعالية تحليل الاقتباسات باستخدام OpenAlex و Scopus، مع التركيز على عدد الاقتباسات، تصنيفات الوثائق، وثلاثة مؤشرات قائمة على الاقتباسات: الأعداد الخام، درجات الاقتباس المعيارية (NCS)، ودرجات الاقتباس المعيارية المحولة لوغاريتميًا (NLCS). قامت الدراسة بتحليل 28.6 مليون مقالة وأجرت 8,704 ارتباطات مقابل معيارين ذهبيين – درجات ChatGPT ومتوسط درجات مراجعة الخبراء من إطار التميز البحثي في المملكة المتحدة (REF) 2021 – تحديدًا لـ 97,816 مقالة. تشير النتائج إلى أن OpenAlex يتفوق على Scopus في عدد الاقتباسات ويحافظ على تصنيف وثائق أكثر شمولاً دون مشاكل كبيرة في تطبيع المجال. من الجدير بالذكر أن الأعداد الخام للاقتباسات كانت فعالة مثل، إن لم تكن أفضل من، المؤشرات المعيارية للتقييمات السنوية الفردية، حيث تفوقت NCS على NLCS.

تشير الاستنتاجات إلى أنه بينما يعد OpenAlex أداة مناسبة لتحليل الاقتباسات، خاصة عندما يكملها تصنيفات Scopus، تتحدى النتائج فائدة تطبيع المجال في بيانات الاقتباسات. تفترض الدراسة أن معدلات الاقتباس الأعلى في بعض المجالات قد ترتبط بأبحاث ذات جودة أعلى، مما يعقد الافتراض بالتماثل عبر التخصصات. يثير هذا مخاوف بشأن صحة المقارنات المباشرة بين المجالات، خاصة في التقييمات متعددة التخصصات. يحذر المؤلفون من أنه بينما قد يسعى مقيمو الأبحاث لتحقيق المساواة المتصورة من خلال تطبيع المجال، قد يؤدي هذا النهج إلى إخفاء الفروق الحقيقية في جودة البحث عبر التخصصات. هناك حاجة لمزيد من التحقيق لتوضيح هذه الديناميات قبل تنفيذ تغييرات في ممارسات تحليل الاقتباسات.

مقدمة

تسلط المقدمة الضوء على أهمية المؤشرات القائمة على الاقتباسات في تقييم الأبحاث عبر كيانات مختلفة، بما في ذلك الأفراد، الأقسام، الجامعات، والدول، كما أشار De Bellis (2009) و Moed (2006). تهدف هذه الدراسة إلى معالجة سؤالين رئيسيين باستخدام نفس مجموعة البيانات: ملاءمة قاعدة بيانات OpenAlex لتحليل الاقتباسات وتحديد المؤشر الأكثر فعالية القائم على الاقتباسات. من المتوقع أن تسهم النتائج من هذه التحقيقات في النقاش المستمر حول موثوقية مقاييس الاقتباس في تقييم الأبحاث الأكاديمية.

الطرق

في هذه الدراسة، استخدم الباحثون نهجًا منهجيًا لتحليل مقالات المجلات من Scopus و OpenAlex للسنوات 2014-2020. تضمنت المنهجية عدة خطوات رئيسية: أولاً، قاموا باستخراج المقالات من كلا القاعدتين ومطابقتها بواسطة DOI لإنشاء ثلاث مجموعات متميزة: مقالات حصرية لـ Scopus، تلك الحصرية لـ OpenAlex، وتلك الموجودة في كلاهما. سمح هذا التقاطع بتحليل مركّز على أسئلة البحث (RQ1، RQ2، RQ3) التي لم تتضمن تباينات في التغطية.

بعد ذلك، قام الباحثون بحساب درجات الاقتباس المعيارية (NLCS) ودرجات الاقتباس غير المعيارية (NCS) لجميع المقالات عبر خمسة مخططات تصنيف ضمن المجموعات الثلاث، مما أسفر عن إجمالي 32 مؤشرًا لكل مقالة. لتأسيس معيار جودة ذهبي، استخدموا ChatGPT لتقييم مقالات REF2021، إلى جانب معيار ثانوي مستمد من متوسط درجات الجودة للقسم. كان الهدف من ارتباط هذه الـ 32 مؤشرًا مع درجات جودة ChatGPT حسب وحدة التقييم (UoA) هو تحديد النهج الأكثر فعالية وتقييم تأثير العوامل المختلفة. من الجدير بالذكر أن الاعتماد الأولي على متوسط درجات REF2021 للقسم أعيد النظر فيه بسبب التحيز العمري الذي يفضل المقالات الأقدم، مما أدى إلى تفضيل درجات ChatGPT، التي أظهرت تحيزًا عمريًا ضئيلًا وقدمت تقييمًا أكثر مباشرة، على الرغم من القيود الجوهرية للحكم غير البشري.

النتائج

حدد التحليل إجمالي 30,123,865 مقالة تحمل معرّفات كائنات رقمية (DOIs) من Scopus و OpenAlex للفترة من 2014 إلى 2020. من الجدير بالذكر أنه بينما كانت جميع المقالات في Scopus مصنفة حسب الموضوع، كانت 1,472,420 مقالة من OpenAlex تفتقر إلى مثل هذه التصنيفات، مما أدى إلى حجم عينة فعالة قدره 28,651,445 مقالة مجلة. علاوة على ذلك، وُجد أن OpenAlex يحتوي على أكثر من ضعف عدد المقالات مقارنة بـ Scopus، كما هو موضح في الجدول 2.

بالإضافة إلى ذلك، كانت مجموعة البيانات المعيارية الذهبية تتكون من 97,816 مقالة مجلة REF2021 التي تحتوي على ملخصات غير قصيرة، كما هو موضح في الجدول 3. من المهم ملاحظة أن بعض المقالات كانت مصنفة تحت وحدات تقييم متعددة (UoAs)، مما أدى إلى إجمالي عدد وحدات التقييم التي تتجاوز عدد المقالات في مجموعة البيانات المجمعة.

المناقشة

تسلط قسم المناقشة في ورقة البحث الضوء على إمكانية OpenAlex كبديل مجاني لفهارس الاقتباسات التقليدية مثل Scopus و Web of Science، خاصة في تعزيز تحليل الاقتباسات للجنوب العالمي وتوفير تغطية أوسع للكتب. ومع ذلك، فإنه يبرز أن التغطية الأكبر لا تعني تلقائيًا تحسين الفائدة لتقييم الأبحاث، حيث أن عوامل مثل تصنيف نوع الوثيقة، تصنيفات الموضوع، ودقة استخراج الاقتباسات تعتبر حاسمة. تشمل توسعات OpenAlex بشكل أساسي المجلات ذات الوصول المفتوح والبيانات، لكنها تواجه تحديات مع دقة البيانات الوصفية، بما في ذلك التصنيف الخاطئ للمساهمات التحريرية والصعوبات في تصنيف اللغة.

تناقش الورقة أيضًا تعقيدات المؤشرات القائمة على الاقتباسات المستخدمة لتقييم جودة البحث. تنتقد الممارسة الشائعة لمقارنة الأعداد الخام للاقتباسات عبر مجالات وأعمار مختلفة، داعيةً إلى استخدام مقاييس معيارية مثل درجة الاقتباس المعيارية (NCS) ودرجة الاقتباس المعيارية المحولة لوغاريتميًا (NLCS) لأخذ هذه الفروقات في الاعتبار. على الرغم من المزايا النظرية لتطبيع المجال، تشير الورقة إلى نقص المقارنات المنهجية ضد معيار ذهبي للتحقق من فعالية مؤشرات الاقتباس المختلفة. تهدف الأبحاث إلى تقييم مصداقية OpenAlex من خلال مقارنة مؤشرات الاقتباس الخاصة بها مع تلك من Scopus وتحديد الصيغ الأكثر فعالية للاقتباس ومخططات تصنيف الموضوع، مما يثير في النهاية تساؤلات حول ضرورة وفعالية تطبيع المجال في تحليل الاقتباسات.

القيود

تسلط القيود في هذه الدراسة الضوء على عدة قيود رئيسية تؤثر على النتائج. أولاً، يثير الاعتماد على تحميل محدد من OpenAlex مخاوف بشأن الطبيعة المتطورة لمحتوياته، وتكنولوجيا الفهرسة، والتصنيفات. كما أن التركيز على المقالات التي تحمل معرّفات كائنات رقمية (DOIs) يعد مشكلة، حيث قد يستبعد العديد من الوثائق ذات الجودة المنخفضة التي تفتقر إلى DOIs. قد يؤدي ذلك إلى تمثيل غير متساوٍ عبر المجالات، مما قد يشوه المؤشرات المعيارية ويضخم بشكل مصطنع الدرجات للمجالات التي تتمتع بانتشار أعلى لمثل هذه الوثائق.

بالإضافة إلى ذلك، فإن التحيزات الموجودة في درجات ChatGPT وارتباطها المحتمل بعدد الاقتباسات تضيف مزيدًا من عدم اليقين. تعتبر درجات REF المستخدمة خشنة، مما يخفف من الارتباطات الملاحظة في المعيار الذهبي الثانوي. من الجدير بالذكر أن درجات جودة ChatGPT تظهر تحيزات ميدانية، تفضل مجالات البحث الأكثر اقتباسًا. بينما لا تؤثر هذه التحيزات بين وحدات التحليل (UoA) مباشرة على الارتباطات داخل UoA المبلغ عنها، فإنها تشير إلى إمكانية وجود تحيزات إضافية لا تزال غير مقدرة. علاوة على ذلك، فإن أعداد الاقتباسات المستخدمة ناضجة نسبيًا، تعكس في الغالب مقالات عمرها خمس سنوات على الأقل، مما قد يشوه ديناميات الارتباط، خاصة بالنسبة للمنشورات الأحدث. وبالتالي، يُحذر من استخلاص استنتاجات قاطعة حول القوة النسبية لأعداد الاقتباسات، درجات الاقتباس المعيارية (NCS)، ودرجات الاقتباس المعيارية المحولة لوغاريتميًا (NLCS) بسبب هذه القيود.

Journal: Journal of the Association for Information Science and Technology, Volume: 76, Issue: 12
DOI: https://doi.org/10.1002/asi.70020
Publication Date: 2025-08-22
Author(s): Mike Thelwall et al.
Primary Topic: scientometrics and bibliometrics research

Overview

This research paper section evaluates the efficacy of citation analysis using OpenAlex and Scopus, focusing on citation counts, document classifications, and three citation-based indicators: raw counts, Normalised Citation Scores (NCS), and Normalised Log-transformed Citation Scores (NLCS). The study analyzed 28.6 million articles and conducted 8,704 correlations against two gold standards—ChatGPT scores and average expert review scores from the UK Research Excellence Framework (REF) 2021—specifically for 97,816 articles. The findings indicate that OpenAlex outperforms Scopus in citation counts and maintains a more inclusive document classification without significant field normalization issues. Notably, raw citation counts were found to be as effective as, if not superior to, normalized indicators for single-year assessments, with NCS outperforming NLCS.

The conclusions suggest that while OpenAlex is a suitable tool for citation analysis, particularly when complemented by Scopus’s classifications, the results challenge the utility of field normalization in citation data. The study posits that higher citation rates in certain fields may correlate with higher quality research, complicating the assumption of equality across disciplines. This raises concerns about the validity of direct comparisons between fields, especially in multidisciplinary evaluations. The authors caution that while research evaluators may strive for perceived equality through field normalization, this approach may obscure genuine differences in research quality across specialties. Further investigation is warranted to clarify these dynamics before implementing changes in citation analysis practices.

Introduction

The introduction highlights the significance of citation-based indicators in evaluating research across various entities, including individuals, departments, universities, and countries, as noted by De Bellis (2009) and Moed (2006). This study aims to address two key questions using the same dataset: the suitability of the OpenAlex database for citation analysis and the identification of the most effective citation-based indicator. The findings from this investigation are expected to contribute to the ongoing discourse on the reliability of citation metrics in academic research evaluation.

Methods

In this study, the researchers employed a systematic approach to analyze journal articles from Scopus and OpenAlex for the years 2014-2020. The methodology involved several key steps: first, they extracted articles from both databases and matched them by DOI to create three distinct sets: articles exclusive to Scopus, those exclusive to OpenAlex, and those present in both. This intersection allowed for focused analysis on research questions (RQ1, RQ2, RQ3) that did not involve coverage discrepancies.

Subsequently, the researchers calculated normalized citation scores (NLCS) and non-normalized citation scores (NCS) for all articles across five classification schemes within the three sets, resulting in a total of 32 indicators per article. To establish a quality gold standard, they utilized ChatGPT to evaluate REF2021 articles, alongside a secondary standard derived from departmental average quality scores. The correlation of these 32 indicators with ChatGPT quality scores by Unit of Assessment (UoA) aimed to determine the most effective approach and assess the impact of various factors. Notably, the initial reliance on departmental mean REF2021 scores was reconsidered due to age bias favoring older articles, leading to the preference for ChatGPT scores, which exhibited minimal age bias and provided a more direct assessment, despite the inherent limitations of non-human judgment.

Results

The analysis identified a total of 30,123,865 articles with Digital Object Identifiers (DOIs) from Scopus and OpenAlex for the period from 2014 to 2020. Notably, while all articles in Scopus were classified by subject, 1,472,420 articles from OpenAlex lacked such classifications, resulting in an effective sample size of 28,651,445 journal articles. Furthermore, OpenAlex was found to contain more than twice the number of articles compared to Scopus, as detailed in Table 2.

Additionally, the gold standard dataset comprised 97,816 REF2021 journal articles that featured non-short abstracts, as presented in Table 3. It is important to note that some articles were categorized under multiple Units of Assessment (UoAs), leading to a total count of UoAs that exceeds the number of articles in the combined dataset.

Discussion

The discussion section of the research paper highlights the potential of OpenAlex as a free alternative to traditional citation indexes like Scopus and Web of Science, particularly in enhancing citation analysis for the Global South and providing broader coverage of books. However, it emphasizes that greater coverage does not automatically translate to improved utility for research evaluation, as factors such as document type classification, subject classifications, and citation extraction accuracy are crucial. OpenAlex’s expansion primarily includes open access and data journals, but it faces challenges with metadata accuracy, including misclassification of editorial contributions and difficulties in language classification.

The paper also addresses the complexities of citation-based indicators used for assessing research quality. It critiques the common practice of comparing raw citation counts across different fields and ages, advocating for normalized metrics like the Normalised Citation Score (NCS) and Normalised Log-transformed Citation Score (NLCS) to account for these disparities. Despite the theoretical advantages of field normalization, the paper notes a lack of systematic comparisons against a gold standard to validate the effectiveness of various citation indicators. The research aims to evaluate OpenAlex’s credibility by comparing its citation indicators with those from Scopus and to identify the most effective citation formulae and subject classification schemes, ultimately questioning the necessity and efficacy of field normalization in citation analysis.

Limitations

The limitations of this study highlight several key constraints affecting the results. Firstly, the reliance on a specific download of OpenAlex raises concerns regarding the evolving nature of its contents, indexing technology, and classifications. The focus on articles with Digital Object Identifiers (DOIs) is also problematic, as it may exclude numerous low-quality documents that lack DOIs. This could lead to an unequal representation across fields, potentially skewing normalized indicators and artificially inflating scores for fields with a higher prevalence of such documents.

Additionally, the biases inherent in the ChatGPT scores and their potential correlation with citation counts introduce further uncertainty. The REF scores utilized are coarse-grained, which dampens the correlations observed in the secondary gold standard. Notably, the ChatGPT quality scores exhibit field biases, favoring more highly cited areas of research. While these between-unit of analysis (UoA) biases do not directly affect the within-UoA correlations reported, they suggest the possibility of additional biases that remain unquantified. Furthermore, the citation counts employed are relatively mature, predominantly reflecting articles that are at least five years old, which may distort the correlation dynamics, particularly for newer publications. Thus, drawing definitive conclusions about the comparative strengths of citation counts, normalized citation scores (NCS), and normalized log citation scores (NLCS) is cautioned against due to these limitations.

شارك: