DOI: https://doi.org/10.24072/pcjournal.687
تاريخ النشر: 2026-03-13
المؤلف: Joel L. Pick وآخرون
الموضوع الرئيسي: الحوسبة العلمية وإدارة البيانات
نظرة عامة
تناقش هذه الفقرة أهمية البيانات المفتوحة والشيفرة في تعزيز الشفافية وقابلية التكرار في البحث العلمي، لا سيما في مجالات علم البيئة وعلم الأحياء التطوري. على الرغم من الاتجاه المتزايد للمجلات التي تتطلب أرشفة البيانات والشيفرة جنبًا إلى جنب مع المقالات المنشورة، إلا أن جودة وكمية هذه الأرشيفات لا تزال غير كافية، مما يؤدي إلى مخاوف بشأن قابلية التكرار. لمعالجة هذه المشكلة، قامت بعض المجلات بتعيين محررين بيانات مخصصين لمساعدة المؤلفين في تحسين جودة بياناتهم وشيفرتهم المؤرشفة. ومع ذلك، هناك نقص في التوافق بشأن المسؤوليات والإجراءات المحددة لمحرري البيانات، مما يؤدي غالبًا إلى عملية غامضة وغير شفافة.
استجابةً لذلك، يقترح المؤلفون، بمشاركة من مختلف محرري البيانات، إرشادات موحدة لمراقبة جودة البيانات والشيفرة تحت رعاية جمعية العلوم المفتوحة والموثوقة والشفافة في علم البيئة وعلم الأحياء التطوري (SORTEE). تقدم هذه الإرشادات إطار عمل مرن من ست مراحل يمكن للمجلات اعتماده تدريجيًا أو على أساس كل حالة على حدة، مع مراعاة الحالات التي قد لا تكون فيها بعض الفحوصات، مثل قابلية التكرار الحسابي، ممكنة. ويخلص المؤلفون إلى أن تنفيذ هذه الإرشادات سيعزز من اتساق وشفافية عملية تحرير البيانات، مما يعود بالنفع على المجلات والمؤلفين والمجتمع العلمي الأوسع من خلال تحسين جودة البيانات والشيفرة المؤرشفة في النهاية.
مقدمة
تسلط المقدمة الضوء على التركيز الكبير على مبادرات العلوم المفتوحة، لا سيما في مجالات علم البيئة وعلم الأحياء التطوري، على مدى العقدين الماضيين. وتؤكد على أهمية البيانات المفتوحة والشيفرة المفتوحة، التي تتضمن الأرشفة العامة لبيانات البحث والشيفرة المرتبطة بها. وقد تم مناقشة الفوائد والأهداف الاجتماعية لمثل هذه الممارسات بشكل موسع في الأدبيات، مع الإشارة إلى أن العديد من الدراسات تشير إلى دعم عام بين الأكاديميين لأرشفة البيانات، حيث يعتقدون أن الفوائد تفوق العيوب المحتملة. ومن الجدير بالذكر أن جمعية العلوم المفتوحة والموثوقة والشفافة في علم البيئة وعلم الأحياء التطوري (SORTEE) قد حددت البيانات المفتوحة والشيفرة المفتوحة كقضايا حاسمة بين أعضائها من 2021 إلى 2025.
علاوة على ذلك، تشير المقدمة إلى أنه استجابةً للطلب المتزايد على البيانات المفتوحة، بدأت العديد من المجلات في هذا المجال في تنفيذ سياسات لتعزيز مشاركة البيانات والشيفرة منذ عام 2010. يعكس هذا التحول التزامًا أوسع لتعزيز الشفافية وقابلية التكرار في البحث العلمي، بما يتماشى مع قيم حركة العلوم المفتوحة.
النتائج
ت outlines قسم النتائج الإرشادات لضمان قابلية التكرار الحسابي لنتائج البحث من خلال أرشفة الشيفرة والبيانات. ويؤكد على أنه يجب تضمين جميع الشيفرات المستخدمة في تصفية البيانات ومعالجتها ورسمها وتحليلها، وأن هيكلها يجب أن يتماشى مع أوصاف المخطوطة. يُكلف محرر البيانات بالتحقق من أن النتائج المقدمة في المخطوطة، بما في ذلك النصوص والجداول والأشكال، يمكن تكرارها من خلال تنفيذ الشيفرة المؤرشفة دون أخطاء. يُتوقع تحقيق قابلية التكرار الدقيقة، ولكن قد تظهر اختلافات نتيجة الطرق العشوائية أو استخدام نماذج اللغة الكبيرة (LLMs)، التي يمكن أن تُدخل عدم اتساق في المخرجات.
يناقش القسم أيضًا تداعيات استخدام نماذج اللغة الكبيرة والبرمجيات المملوكة، مشيرًا إلى أنه بينما يمكن أرشفة الشيفرة، قد لا تكون النتائج قابلة للتكرار بسبب التباين الفطري. يقترح أن يتم تقييم الانحرافات عن النتائج المتوقعة باستخدام نسبة الخطأ، مع تحديد عتبات محددة للانحرافات الطفيفة والكبيرة. يجب على محرر البيانات التواصل بشأن شروط تقييم القابلية للتكرار، بما في ذلك أي عتبات تحمل. إذا لم يكن من الممكن تقييم القابلية للتكرار الحسابي بالكامل، يجب أن يتم توضيح هذا القيد بوضوح في المخطوطة، ويجب تقديم وتوثيق أي مخرجات من الشيفرة غير القابلة للتنفيذ. بشكل عام، تدعو الإرشادات إلى الشفافية والمرونة في عملية تقييم القابلية للتكرار.
المناقشة
ت outlines قسم المناقشة من ورقة البحث الأهداف الحاسمة لأرشفة البيانات والشيفرة، مؤكدًا على أهمية تمكين إعادة استخدام البيانات، وتعزيز الشفافية، وتحقيق القابلية للتكرار الحسابي. يبرز المؤلفون أن أرشفة البيانات تهدف إلى منع فقدان البيانات، وتسهيل تجميع الأبحاث من خلال التحليل التلوي، والسماح بإعادة تحليل البيانات الموجودة باستخدام منهجيات محدثة. ويؤكدون على أن الالتزام بمبادئ FAIR (قابلة للاكتشاف، قابلة للوصول، قابلة للتشغيل المتبادل، قابلة لإعادة الاستخدام) أمر ضروري لتعظيم فائدة البيانات المؤرشفة، على الرغم من أن العديد من مجموعات البيانات الحالية لا تلبي هذه المعايير.
علاوة على ذلك، يبرز القسم دور الشفافية في بناء الثقة داخل المجتمع العلمي، لا سيما في ضوء التعقيد المتزايد في التحليلات والحالات السابقة من الاحتيال الأكاديمي. يوفر الوصول إلى البيانات والشيفرة إمكانية التحقق من النتائج وتحديد الأخطاء المحتملة، مما يعزز من مصداقية الأبحاث المنشورة. كما يعرف المؤلفون القابلية للتكرار الحسابي على أنها القدرة على الحصول على نتائج متسقة باستخدام نفس البيانات والأساليب، مشيرين إلى أن تحقيق هذا الهدف يتطلب اهتمامًا دقيقًا بالتفاصيل في عملية الأرشفة، بما في ذلك توفير شيفرة كاملة وعاملة.
أخيرًا، تناقش الورقة إنشاء محرري بيانات في المجلات للإشراف على مراقبة جودة البيانات والشيفرة، بهدف تحسين الجودة العامة للمواد المؤرشفة. يتم تصوير هذه المبادرة على أنها مفيدة ليس فقط للمؤلفين ولكن أيضًا للمجلات والمجتمع البحثي الأوسع، حيث تعزز من معايير النزاهة العلمية وتسهيل التعاون. يدعو المؤلفون إلى إرشادات موحدة لتبسيط عملية مراقبة جودة البيانات والشيفرة عبر المجلات، مما يعزز في النهاية ثقافة الانفتاح والموثوقية في الأبحاث البيئية والتطورية.
DOI: https://doi.org/10.24072/pcjournal.687
Publication Date: 2026-03-13
Author(s): Joel L. Pick et al.
Primary Topic: Scientific Computing and Data Management
Overview
The section discusses the importance of open data and code in enhancing transparency and reproducibility in scientific research, particularly within the fields of ecology and evolutionary biology. Despite the growing trend of journals requiring data and code to be archived alongside published articles, the quality and quantity of such archives remain insufficient, leading to concerns about reproducibility. To address this issue, some journals have appointed dedicated data editors to assist authors in improving the quality of their archived data and code. However, there is a lack of consensus on the specific responsibilities and procedures for data editors, which often results in a vague and opaque process.
In response, the authors, with input from various data editors, propose standardized guidelines for Data and Code Quality Control under the auspices of the Society for Open, Reliable, and Transparent Ecology and Evolutionary Biology (SORTEE). These guidelines present a flexible six-stage framework that journals can adopt incrementally or on a case-by-case basis, accommodating situations where certain checks, such as computational reproducibility, may not be feasible. The authors conclude that the implementation of these guidelines will enhance the consistency and transparency of the data editing process, benefiting journals, authors, and the broader scientific community by ultimately improving the quality of archived data and code.
Introduction
The introduction highlights the significant emphasis on open science initiatives, particularly in the fields of ecology and evolutionary biology, over the past two decades. It underscores the importance of open data and open code, which involve the public archiving of research data and associated code. The societal benefits and goals of such practices have been extensively discussed in the literature, with various studies indicating a general support among academics for data archiving, as they believe the advantages surpass potential drawbacks. Notably, the Society for Open, Reliable, and Transparent Ecology and Evolutionary Biology (SORTEE) has identified open data and open code as critical issues among its members from 2021 to 2025.
Furthermore, the introduction notes that in response to the growing demand for open data, numerous journals in the field have begun implementing policies to promote data and code sharing since 2010. This shift reflects a broader commitment to enhancing transparency and reproducibility in scientific research, aligning with the values of the open science movement.
Results
The results section outlines the guidelines for ensuring computational reproducibility of research findings through the archiving of code and data. It emphasizes that all code used for data filtering, processing, graphing, and analysis must be included and that its structure should align with the manuscript’s descriptions. The data editor is tasked with verifying that the results presented in the manuscript, including text, tables, and figures, can be reproduced by executing the archived code without errors. Exact reproducibility is expected, but variations may arise from stochastic methods or the use of large language models (LLMs), which can introduce inconsistencies in outputs.
The section further discusses the implications of using LLMs and proprietary software, noting that while the code can be archived, the results may not be reproducible due to inherent variability. It suggests that deviations from expected results should be assessed using percentage error, with specific thresholds for minor and major deviations. The data editor should communicate the conditions of reproducibility assessment, including any tolerance thresholds. If computational reproducibility cannot be fully assessed, this limitation should be clearly stated in the manuscript, and any outputs from non-executable code must be provided and documented. Overall, the guidelines advocate for transparency and flexibility in the reproducibility assessment process.
Discussion
The discussion section of the research paper outlines the critical goals of data and code archiving, emphasizing the importance of enabling data reuse, enhancing transparency, and achieving computational reproducibility. The authors highlight that data archiving serves to prevent data loss, facilitate research synthesis through meta-analysis, and allow for the re-analysis of existing data with updated methodologies. They stress that adherence to the FAIR principles (Findable, Accessible, Interoperable, Reusable) is essential for maximizing the utility of archived data, although many datasets currently fall short of these standards.
Furthermore, the section underscores the role of transparency in building trust within the scientific community, particularly in light of increasing complexity in analyses and past instances of academic fraud. Providing access to data and code allows for the verification of results and the identification of potential errors, thereby enhancing the credibility of published research. The authors also define computational reproducibility as the ability to obtain consistent results using the same data and methods, noting that achieving this goal requires meticulous attention to detail in the archiving process, including the provision of complete and functional code.
Lastly, the paper discusses the establishment of data editors in journals to oversee data and code quality control, aiming to improve the overall quality of archived materials. This initiative is portrayed as beneficial not only for authors but also for journals and the broader research community, as it promotes higher standards of scientific integrity and facilitates collaboration. The authors advocate for standardized guidelines to streamline the data and code quality control process across journals, ultimately fostering a culture of openness and reliability in ecological and evolutionary research.
