DOI: https://doi.org/10.1007/s11914-023-00852-0
PMID: https://pubmed.ncbi.nlm.nih.gov/38227177
تاريخ النشر: 2024-01-16
المؤلف: Melissa A. Kacena وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في الرعاية الصحية والتعليم
نظرة عامة
كان الهدف من هذه المراجعة هو تقييم فعالية ChatGPT في مساعدة كتابة مقالات مراجعة علمية موثوقة، تمت مراجعتها من قبل الأقران، وخاصة في سياق أبحاث الجهاز العضلي الهيكلي. ركزت الدراسة على ثلاثة مواضيع رئيسية: العلاقة بين مرض الزهايمر وصحة العظام، التنظيم العصبي لشفاء الكسور، وتأثير COVID-19 على صحة الجهاز العضلي الهيكلي. تم استخدام ثلاثة أساليب لصياغة المخطوطات: بشرية فقط، ذكاء اصطناعي فقط، ومزيج من الاثنين. خضعت المقالات الناتجة لعملية تدقيق شاملة وتحرير لضمان الصرامة العلمية، مما كشف عن اختلافات كبيرة عن المسودات الأولية.
أشارت النتائج الأخيرة إلى أنه بينما قلل استخدام الذكاء الاصطناعي من الوقت المطلوب لكتابة مقالات المراجعة، إلا أنه أدخل أيضًا أخطاء كبيرة، حيث كانت تصل نسبة المراجع المقتبسة في المسودات التي تعتمد على الذكاء الاصطناعي فقط إلى 70% غير صحيحة. كانت الطريقة المدعومة بالذكاء الاصطناعي، على الرغم من كونها أكثر دقة من الطريقة التي تعتمد على الذكاء الاصطناعي فقط، تثير القلق بشأن الانتحال، حيث أنتجت أعلى مؤشرات تشابه. بالإضافة إلى ذلك، كانت قيود ChatGPT 4.0، الذي كان لديه حد معرفة في سبتمبر 2021، تتطلب إدخال يدوي للأدبيات الأكثر حداثة، مما أدى إلى تقارب بين الأساليب التي تعتمد على الذكاء الاصطناعي فقط والمدعومة بالذكاء الاصطناعي للاقتباسات المعاصرة. في النهاية، خلصت الدراسة إلى أنه بينما يمكن للذكاء الاصطناعي تسريع عملية الكتابة، إلا أنه لا يمكن الاعتماد عليه بشكل مستقل بسبب عدم دقته؛ بدلاً من ذلك، يجب استخدامه تحت إشراف بشري دقيق لتعزيز عملية الكتابة العلمية.
مقدمة
تناقش مقدمة هذه الورقة البحثية الإمكانية التحويلية للذكاء الاصطناعي (AI)، وخاصة نماذج اللغة الذكية مثل ChatGPT، في تعزيز الكفاءة في الكتابة العلمية. مع ظهور المحولات المدربة مسبقًا (GPT) والتقدم في معالجة اللغة الطبيعية (NLP)، بدأ الذكاء الاصطناعي يتخلل مجالات مختلفة، بما في ذلك الطب والبحث. بينما يمكن لأدوات الذكاء الاصطناعي المساعدة في القواعد، والمفردات، والبحث في الأدبيات، لا يزال هناك شك بشأن استخدامها في كتابة المقالات الأكاديمية بسبب المخاوف المتعلقة بالدقة، وانتهاك حقوق الطبع والنشر، وظاهرة “الهلاوس الاصطناعية”، حيث ينتج الذكاء الاصطناعي معلومات مضللة.
يقترح المؤلفون دراسة لتقييم فعالية الذكاء الاصطناعي في كتابة مقالات مراجعة علمية كاملة من خلال مقارنة ثلاث استراتيجيات: (1) نهج تقليدي يعتمد على البشر فقط، (2) طريقة تعتمد على الذكاء الاصطناعي فقط، و(3) نهج مدعوم بالذكاء الاصطناعي حيث يقوم البشر بإنتاج مراجعة الأدبيات والمخطط. يفترضون أن العملية التي تعتمد على البشر فقط ستكون الأكثر استهلاكًا للوقت ولكنها ستتطلب أقل عدد من التعديلات، بينما ستكون الطريقة التي تعتمد على الذكاء الاصطناعي فقط الأسرع ولكنها عرضة للأخطاء. تركز الدراسة على ثلاثة مواضيع في مجال الجهاز العضلي الهيكلي، بهدف تقييم قدرة الذكاء الاصطناعي على إنتاج عمل بجودة قابلة للنشر مع معالجة التحديات والقيود المرتبطة باستخدامه في الكتابة العلمية.
الطرق
تحدد قسم “الطرق” تصميم التجربة والتقنيات التحليلية المستخدمة في الدراسة. استخدم الباحثون نهجًا كميًا، حيث نفذوا تجارب محكومة لجمع البيانات حول المتغيرات المحددة. تم إجراء تحليلات إحصائية باستخدام أدوات برمجية لضمان موثوقية وصدق النتائج، مع تحديد مستويات الدلالة عند p < 0.05. شملت جمع البيانات طريقة أخذ عينات منهجية، مما يضمن عينة تمثيلية من السكان قيد الدراسة. طبق الباحثون نماذج رياضية مختلفة لتفسير البيانات، بما في ذلك تحليل الانحدار لتحديد العلاقات بين المتغيرات المستقلة والتابعة. تم تصميم المنهجية لتقليل التحيز وتعزيز إمكانية التكرار، مما يعزز النتائج العامة للبحث.
المناقشة
في قسم المناقشة من الورقة البحثية، يتم توضيح المنهجيات المستخدمة لكتابة مقالات المراجعة باستخدام الأساليب التي تعتمد على البشر، والذكاء الاصطناعي فقط، والمدعومة بالذكاء الاصطناعي. اتبعت المقالات التي تم إنشاؤها بواسطة البشر الممارسات التقليدية، بما في ذلك مراجعات الأدبيات الشاملة، والمخططات المنظمة، والتحرير التعاوني بين المؤلفين المشاركين. تم تسهيل إدارة المراجع من خلال EndNote، وتم إنشاء ملخصات رسومية باستخدام BioRender.com، حيث تم إجراء عملية الكتابة بالكامل بواسطة مؤلفين بشريين.
على النقيض من ذلك، تم صياغة مقالات المراجعة التي تعتمد على الذكاء الاصطناعي فقط باستخدام نسخة ChatGPT Plus مع نموذج GPT-4، الذي كان محدودًا بحد معرفته في سبتمبر 2021. جرب المؤلفون استفسارات مختلفة لتوليد المخططات والمحتوى، ولكن بسبب الجدول الزمني لنشر المقالات ذات الصلة، غالبًا ما كانت المخطوطات التي تم إنشاؤها بواسطة الذكاء الاصطناعي تتطلب تدخلًا بشريًا للتحقق من الحقائق ودقة الاقتباسات. استخدمت المقالات المدعومة بالذكاء الاصطناعي إطار عمل مشابه ولكنها دمجت المخططات والمراجع التي أنشأها البشر، مستفيدة من مكون AskYourPDF لتحليل الوثائق. سمح هذا النهج بتوليف أكثر تكاملاً للأدبيات الحديثة، على الرغم من أنه تطلب أيضًا إدارة دقيقة للاستفسارات وتحميل الوثائق لضمان تغطية شاملة للمواضيع.
بشكل عام، قيست الدراسة معايير مختلفة عبر ثلاثة أساليب كتابة، بما في ذلك الوقت المستغرق في أنشطة مختلفة، ودرجات التشابه بين المسودات، ودقة المراجع. كانت هذه التقييمات تهدف إلى تقييم كفاءة وفعالية كل طريقة في إنتاج مقالات مراجعة علمية صارمة، مما يبرز الفوائد والقيود المحتملة لدمج الذكاء الاصطناعي في الكتابة الأكاديمية.
DOI: https://doi.org/10.1007/s11914-023-00852-0
PMID: https://pubmed.ncbi.nlm.nih.gov/38227177
Publication Date: 2024-01-16
Author(s): Melissa A. Kacena et al.
Primary Topic: Artificial Intelligence in Healthcare and Education
Overview
The purpose of this review was to evaluate the efficacy of ChatGPT in assisting the writing of credible, peer-reviewed scientific review articles, particularly in the context of musculoskeletal research. The study focused on three key topics: the relationship between Alzheimer’s disease and bone health, neural regulation of fracture healing, and the impact of COVID-19 on musculoskeletal health. Three manuscript drafting approaches were employed: human-only, AI-only, and a combination of both. The resulting articles underwent extensive fact-checking and editing to ensure scientific rigor, revealing significant differences from the initial drafts.
Recent findings indicated that while the use of AI reduced the time required to write the review articles, it also introduced substantial inaccuracies, with up to 70% of references cited in the AI-only drafts being incorrect. The AI-assisted approach, although more accurate than the AI-only method, raised concerns regarding plagiarism, as it produced the highest similarity indices. Additionally, the limitations of ChatGPT 4.0, which had a knowledge cutoff in September 2021, necessitated manual input of more recent literature, leading to a convergence of the AI-only and AI-assisted approaches for contemporary citations. Ultimately, the study concluded that while AI can expedite the writing process, it cannot be relied upon independently due to its inaccuracies; instead, it should be utilized under careful human oversight to enhance the scientific writing process.
Introduction
The introduction of this research paper discusses the transformative potential of artificial intelligence (AI), particularly AI language models like ChatGPT, in enhancing efficiency in scientific writing. With the advent of generative pretrained transformers (GPT) and advancements in natural language processing (NLP), AI has begun to permeate various fields, including medicine and research. While AI tools can assist with grammar, vocabulary, and literature searches, there remains skepticism regarding their use in writing scholarly articles due to concerns about accuracy, copyright infringement, and the phenomenon of “artificial hallucinations,” where AI generates misleading information.
The authors propose a study to evaluate the effectiveness of AI in writing full-length scientific review articles by comparing three strategies: (1) a traditional human-only approach, (2) an AI-only method, and (3) an AI-assisted approach where humans generate the literature review and outline. They hypothesize that the human-only process will be the most time-consuming but require the fewest revisions, while the AI-only method will be the fastest but prone to inaccuracies. The study focuses on three topics in the musculoskeletal field, aiming to assess AI’s ability to produce publishable quality work while addressing the challenges and limitations associated with its use in scientific writing.
Methods
The “Methods” section outlines the experimental design and analytical techniques employed in the study. The researchers utilized a quantitative approach, implementing controlled experiments to gather data on the specified variables. Statistical analyses were conducted using software tools to ensure the reliability and validity of the results, with significance levels set at p < 0.05. Data collection involved a systematic sampling method, ensuring a representative sample of the population under study. The researchers applied various mathematical models to interpret the data, including regression analysis to identify relationships between independent and dependent variables. The methodology was designed to minimize bias and enhance reproducibility, thereby strengthening the overall findings of the research.
Discussion
In the discussion section of the research paper, the methodologies employed for writing review articles using human-generated, AI-only, and AI-assisted approaches are delineated. The human-generated articles followed traditional practices, including comprehensive literature reviews, structured outlines, and collaborative editing among co-authors. Reference management was facilitated through EndNote, and graphical abstracts were created using BioRender.com, with the entire writing process being conducted by human authors.
Conversely, the AI-only review articles were drafted using the ChatGPT Plus version with the GPT-4 model, which was limited by its knowledge cutoff in September 2021. The authors experimented with various queries to generate outlines and content, but due to the publication timeline of relevant articles, the AI-generated manuscripts often required human intervention for fact-checking and citation accuracy. The AI-assisted articles utilized a similar framework but incorporated human-generated outlines and references, leveraging the AskYourPDF plugin for document analysis. This approach allowed for a more integrated synthesis of recent literature, although it also necessitated careful management of queries and document uploads to ensure comprehensive coverage of the topics.
Overall, the study measured various parameters across the three writing approaches, including time spent on different activities, similarity scores between drafts, and the accuracy of references. These assessments aimed to evaluate the efficiency and effectiveness of each method in producing scientifically rigorous review articles, highlighting the potential benefits and limitations of AI integration in academic writing.
