DOI: https://doi.org/10.3389/frai.2024.1341697
PMID: https://pubmed.ncbi.nlm.nih.gov/38384276
تاريخ النشر: 2024-02-07
المؤلف: Dorian Quelle وآخرون
الموضوع الرئيسي: المعلومات المضللة وتأثيراتها
نظرة عامة
تقيّم ورقة البحث فعالية GPT-3.5 و GPT-4 من OpenAI في التحقق الآلي من الحقائق، مع تسليط الضوء على الحاجة المتزايدة لأدوات التعلم الآلي في مكافحة المعلومات المضللة. تكشف الدراسة أنه بينما تظهر كلا النموذجين دقة جديرة بالثناء – 63-75% في المتوسط دون سياق، تتحسن إلى أكثر من 80% و 89% للمطالب غير الغامضة مع السياق – يتفوق GPT-4 باستمرار على GPT-3.5. ومن الجدير بالذكر أن لغة الطلب تؤثر بشكل كبير على الأداء، حيث غالبًا ما تؤدي الطلبات المترجمة إلى نتائج أفضل من نظيراتها الأصلية، خاصة في السياقات غير الإنجليزية. وهذا يبرز اعتماد النماذج على اللغة الأصلية للمطالب وآثار تمثيل البيانات في مجموعات بيانات التدريب.
تناقش الورقة أيضًا المخاوف المتعلقة بتسرب البيانات، مشيرة إلى أنه بينما تميل المعلومات المضللة إلى التكرار عبر المنصات واللغات، فإن هذه الظاهرة لا تعيق أداء النماذج في التحقق من المطالب الجديدة. في الواقع، يبدو أن القدرة على دمج السياق في الوقت الحقيقي تعزز من قدرات التفكير. ومع ذلك، تشير التباينات في الأداء عبر لغات مختلفة إلى أن التمثيل المنحرف للبيانات التدريبية، الذي يفضل اللغة الإنجليزية بشكل خاص، يؤثر على الدقة. وبالتالي، تشير النتائج إلى أنه بينما يمكن لنماذج اللغة الكبيرة (LLMs) المساعدة في التحقق من الحقائق، إلا أنها لا يمكن أن تحل محل المدققين البشريين بالكامل بسبب الأخطاء العرضية وتعقيدات المعلومات المضللة متعددة اللغات. هناك حاجة لمزيد من البحث لتعميق الفهم حول نقاط القوة والقيود للنماذج في سياقات لغوية متنوعة.
مقدمة
تؤكد مقدمة ورقة البحث هذه على الدور الحاسم للتحقق من الحقائق في مكافحة المعلومات المضللة، خاصة في سياق الأحداث السياسية الأخيرة وجائحة COVID-19. تسلط الضوء على تطور منظمات التحقق من الحقائق، ولا سيما تأسيس Politifact في عام 2009، الذي قدم نظام تصنيف ترتيبي أثر على المبادرات العالمية للتحقق من الحقائق. تشير الورقة إلى أن ارتفاع المعلومات المضللة، الذي تفاقم بسبب الاستقطاب السياسي وانتشار وسائل التواصل الاجتماعي، يتطلب طرق تحقق من الحقائق أكثر كفاءة، حيث أن العمليات اليدوية التقليدية أصبحت غير كافية بشكل متزايد.
يناقش المؤلفون إمكانيات نماذج اللغة الكبيرة (LLMs) مثل GPT-4 في تعزيز كفاءة التحقق من الحقائق، مع الاعتراف أيضًا بالتحديات الكبيرة، مثل اعتماد النماذج على بيانات التدريب التي قد تشمل التحقق من الحقائق نفسها. يثير هذا مخاوف بشأن قدرة النماذج على التعميم خارج أمثلة التدريب الخاصة بها. توضح الورقة مجموعة متنوعة من الأساليب الآلية للتحقق من الحقائق والمهام المشتركة التي تهدف إلى تحسين دقة وشفافية هذه الأنظمة. يقترح المؤلفون تقييمًا جديدًا لـ GPT-3.5 و GPT-4، مع التركيز على أدائهم في التحقق من الحقائق مع وبدون سياق خارجي، ويقدمون منهجية تدمج استرجاع المعلومات مع التحقق من المطالب. تشير نتائجهم إلى أن المعلومات السياقية تعزز الدقة بشكل كبير، حيث يتفوق GPT-4 على GPT-3.5، خاصة عندما يتم ترجمة المطالب إلى الإنجليزية. بشكل عام، تؤكد الدراسة على أهمية تطوير أنظمة آلية قوية لمواجهة التحديات التي تطرحها المعلومات المضللة في مشهد المعلومات اليوم.
الطرق
تحدد ورقة البحث نهجًا جديدًا للتحقق الآلي من الحقائق يدمج استرجاع الأدلة والتحقق من المطالب باستخدام نماذج اللغة الكبيرة (LLMs)، وبشكل خاص GPT-3.5 و GPT-4. تتضمن المنهجية شرطين أساسيين لتقييم أداء النماذج: يعتمد الشرط الأول فقط على المعرفة السابقة للنماذج بالمطالب، بينما يسمح الشرط الثاني للنماذج بالاستعلام عن معلومات سياقية ذات صلة من بحث Google. يتيح دمج إطار Reasoning and Acting (ReAct) لنماذج اللغة الكبيرة التفاعل مع الأدوات الخارجية، مما يسهل عملية ديناميكية حيث يمكن للنماذج أن تقرر استرجاع معلومات إضافية بناءً على ملاحظاتها.
لتحسين عملية استرجاع المعلومات، جربت الدراسة في البداية محتوى HTML الكامل من المواقع ولكن وجدت أنه مرهق لنافذة سياق LLM. بدلاً من ذلك، تم استخدام وظيفة استرجاع المعلومات BM25 لتقطير المحتوى الحيوي، على الرغم من عدم ملاحظة أي تحسين في الأداء، حيث أن طرق التعلم الآلي الحالية من Google تلخص بالفعل المعلومات المهمة بشكل فعال. تم تجهيز نماذج اللغة الكبيرة لاسترجاع ما يصل إلى 10 نتائج بحث لكل تكرار، وإذا فشلت في تقديم إجابة حاسمة بعد ثلاث محاولات، يتم تحفيزها لتلخيص نتائجها بناءً على المعلومات التي تم جمعها. تشير الورقة أيضًا إلى استخدام مكتبة LangChain لتنفيذ الوكيل داخل إطار ReAct، مما يعرض قدرات النظام من خلال أمثلة للمطالب التي تم تقييمها بشكل صحيح وغير صحيح.
النتائج
تكشف نتائج الدراسة حول التحقق الآلي من الحقائق باستخدام مجموعة بيانات PolitiFact عن اختلافات كبيرة في الأداء بين نماذج GPT-3.5 و GPT-4. كما هو موضح في الجدول 1، عند تقييمها في ظل شرط عدم وجود سياق، يتفوق GPT-4 عمومًا على GPT-3.5، خاصة في التنبؤ بالمطالب الكاذبة، حيث حدد GPT-3.5 58.2% من المطالب على أنها كاذبة مقارنة بـ 22.89% لـ GPT-4. ومع ذلك، يظهر GPT-3.5 أداءً متفوقًا في فئة “كاذب إلى حد كبير”، بينما تحقق كلا النموذجين دقة تزيد عن 90% للمطالب المصنفة على أنها “كاذب بشكل واضح”.
في المقابل، تحت شرط السياق، يظهر GPT-3.5 معايرة أفضل، حيث يحقق دقة أكثر توازنًا بين الأحكام الصحيحة والخاطئة. ومن الجدير بالذكر أن GPT-3.5 يظهر أداءً محسنًا في التنبؤ بالأحكام الصحيحة مقارنة بشرط عدم وجود سياق، بينما يستفيد GPT-4 أيضًا من السياق، مما يزيد من دقته في الفئات الصحيحة بمعدل 10.19 نقطة مئوية. بشكل عام، يؤدي كلا النموذجين أداءً أفضل في شرط السياق، مما يعزز قدرتهما على التمييز بين المطالب الصحيحة والخاطئة ويحسن من معايرتهما في التنبؤ بالنتائج الصحيحة.
المناقشة
في هذه الدراسة، أجرى المؤلفون تجارب لتقييم قدرات التحقق من الحقائق لنماذج GPT-3.5 و GPT-4 من OpenAI. استخدمت التجربة الأولى مجموعة بيانات PolitiFact، التي تحتوي على 21,152 مطالبة تم التحقق منها مصنفة إلى ستة مستويات من الحقيقة. أشارت النتائج إلى أن GPT-4 تفوق بشكل كبير على GPT-3.5، خاصة عندما تم توفير معلومات سياقية، حيث حقق معدلات دقة تتجاوز 80% للأحكام غير الغامضة. أظهرت النماذج أداءً أفضل في تحديد البيانات الكاذبة مقارنة بالبيانات الصحيحة، مما يبرز التحديات الكامنة في تصنيف المطالب الدقيقة مثل “نصف صحيح” و “كاذب إلى حد كبير”.
ركزت التجربة الثانية على مجموعة بيانات متعددة اللغات من Data Commons، وكشفت أن كلا النموذجين واجها صعوبة مع المطالب غير الإنجليزية، على الرغم من أن الطلبات المترجمة عمومًا حققت نتائج أفضل. أظهرت التحليلات أن دقة النماذج تختلف بشكل كبير عبر اللغات، حيث تتفوق الترجمات الإنجليزية باستمرار على الطلبات بلغة الأصل. ومن الجدير بالذكر أن الدراسة لم تجد أي تراجع في الأداء للتحقق من الحقائق الذي تم بعد تاريخ قطع تدريب النماذج، مما يشير إلى أن استرجاع السياق في الوقت الحقيقي قد يعزز من قدرات التفكير للنماذج فيما يتعلق بالمعلومات المضللة الجديدة. بشكل عام، بينما تؤكد النتائج على إمكانيات نماذج اللغة الكبيرة في دعم جهود التحقق من الحقائق، فإنها تؤكد أيضًا على ضرورة الإشراف البشري بسبب قيود النماذج في التمييز بدقة بين صحة المطالب المعقدة عبر لغات متنوعة.
DOI: https://doi.org/10.3389/frai.2024.1341697
PMID: https://pubmed.ncbi.nlm.nih.gov/38384276
Publication Date: 2024-02-07
Author(s): Dorian Quelle et al.
Primary Topic: Misinformation and Its Impacts
Overview
The research paper evaluates the efficacy of OpenAI’s GPT-3.5 and GPT-4 in automated fact-checking, highlighting the increasing necessity of machine learning tools in combating misinformation. The study reveals that while both models demonstrate commendable accuracy—63-75% on average without context, improving to over 80% and 89% for non-ambiguous claims with context—GPT-4 consistently outperforms GPT-3.5. Notably, the language of the prompt significantly influences performance, with translated prompts often yielding better results than their original counterparts, particularly in non-English contexts. This underscores the models’ dependency on the original language of claims and the implications of data representation in training datasets.
The paper also addresses concerns regarding data leakage, suggesting that while misinformation tends to recur across platforms and languages, this phenomenon does not hinder the models’ performance in fact-checking novel claims. In fact, the ability to incorporate real-time context appears to enhance reasoning capabilities. However, the variability in performance across different languages indicates that the training data’s skewed representation, particularly favoring English, affects accuracy. Consequently, the findings suggest that while LLMs can assist in fact-checking, they cannot fully replace human fact-checkers due to occasional inaccuracies and the complexities of multilingual misinformation. Further research is warranted to deepen understanding of the models’ strengths and limitations in diverse linguistic contexts.
Introduction
The introduction of this research paper emphasizes the critical role of fact-checking in combating misinformation, particularly in the context of recent political events and the COVID-19 pandemic. It highlights the evolution of fact-checking organizations, notably the establishment of Politifact in 2009, which introduced an ordinal rating system that has influenced global fact-checking initiatives. The paper notes that the rise of misinformation, exacerbated by political polarization and the proliferation of social media, necessitates more efficient fact-checking methods, as traditional manual processes are increasingly inadequate.
The authors discuss the potential of large language models (LLMs) like GPT-4 in enhancing fact-checking efficiency, while also acknowledging significant challenges, such as the models’ reliance on training data that may include fact-checks themselves. This raises concerns about the models’ ability to generalize beyond their training examples. The paper outlines various automated fact-checking approaches and shared tasks aimed at improving the accuracy and explainability of these systems. The authors propose a novel evaluation of GPT-3.5 and GPT-4, focusing on their performance in fact-checking with and without external context, and introduce a methodology that integrates information retrieval with claim verification. Their findings indicate that contextual information significantly enhances accuracy, with GPT-4 outperforming GPT-3.5, particularly when claims are translated into English. Overall, the study underscores the importance of developing robust automated systems to address the challenges posed by misinformation in today’s information landscape.
Methods
The research paper outlines a novel approach to automated fact-checking that integrates evidence retrieval and claim verification using large language models (LLMs), specifically GPT-3.5 and GPT-4. The methodology involves two primary conditions for evaluating the models’ performance: the first condition relies solely on the models’ pre-existing knowledge of claims, while the second condition allows the models to query Google Search for relevant contextual information. The integration of the Reasoning and Acting (ReAct) framework enables the LLMs to interact with external tools, facilitating a dynamic process where the models can decide to retrieve additional information based on their observations.
To optimize the information retrieval process, the study initially experimented with full HTML content from websites but found it overwhelming for the LLM’s context window. Instead, the BM25 information retrieval function was employed to distill critical content, although no performance improvement was observed, as Google’s existing machine learning methods already effectively summarize important information. The LLMs are equipped to retrieve up to 10 search results per iteration, and if they fail to provide a conclusive answer after three attempts, they are prompted to summarize their findings based on the gathered information. The paper also notes the use of the LangChain library to implement the agent within the ReAct framework, showcasing the system’s capabilities through examples of correctly and incorrectly assessed claims.
Results
The results of the study on automated fact-checking using the PolitiFact dataset reveal significant differences in performance between the GPT-3.5 and GPT-4 models. As shown in Table 1, when evaluated under a no-context condition, GPT-4 generally outperforms GPT-3.5, particularly in predicting false claims, with GPT-3.5 identifying 58.2% of claims as false compared to 22.89% for GPT-4. However, GPT-3.5 demonstrates superior performance in the “mostly false” category, while both models achieve over 90% accuracy for claims labeled “pants-on-fire.”
In contrast, under a context condition, GPT-3.5 exhibits better calibration, achieving a more balanced accuracy between true and false verdicts. Notably, GPT-3.5 shows improved performance in predicting true verdicts compared to the no-context condition, while GPT-4 also benefits from context, increasing its accuracy in true categories by an average of 10.19 percentage points. Overall, both models perform better in the context condition, enhancing their ability to distinguish between true and false claims and improving their calibration in predicting true outcomes.
Discussion
In this study, the authors conducted experiments to evaluate the fact-checking capabilities of OpenAI’s GPT-3.5 and GPT-4 models. The first experiment utilized the PolitiFact dataset, which contains 21,152 fact-checked claims categorized into six ordinal truth levels. The results indicated that GPT-4 significantly outperformed GPT-3.5, particularly when contextual information was provided, achieving accuracy rates exceeding 80% for non-ambiguous verdicts. The models demonstrated better performance in identifying false statements compared to true ones, highlighting the inherent challenges in classifying nuanced claims such as “half-true” and “mostly-false.”
The second experiment focused on a multilingual dataset from Data Commons, revealing that both models struggled with non-English claims, although translated prompts generally yielded better results. The analysis showed that the accuracy of the models varied significantly across languages, with English translations consistently outperforming original language prompts. Notably, the study found no decline in performance for fact-checks conducted after the models’ training cut-off date, suggesting that real-time context retrieval could enhance the models’ reasoning capabilities regarding novel misinformation. Overall, while the findings underscore the potential of LLMs in assisting fact-checking efforts, they also emphasize the necessity for human oversight due to the models’ limitations in accurately discerning the veracity of complex claims across diverse languages.
