DOI: https://doi.org/10.1038/s41598-025-85308-4
PMID: https://pubmed.ncbi.nlm.nih.gov/39774143
تاريخ النشر: 2025-01-07
المؤلف: Jun Wang وآخرون
الموضوع الرئيسي: المعلومات المضللة وتأثيراتها
نظرة عامة
تتناول الدراسة القضية الملحة لانتشار المعلومات المضللة في الشبكات الاجتماعية المحمولة، والتي تشكل مخاطر على الاستقرار الاجتماعي، والصحة العامة، والعمليات الديمقراطية. تقدم نموذجًا هجينًا يدمج تمثيلات الترميز ثنائية الاتجاه من المحولات (BERT) مع الشبكات العصبية الذاكرة القصيرة والطويلة (LSTM)، بهدف تعزيز الكشف المبكر عن المعلومات المضللة بناءً فقط على المحتوى النصي. أظهر نموذج BERT-LSTM مقاييس أداء مثيرة للإعجاب، حيث حقق دقة بنسبة 93.51%، واسترجاع بنسبة 91.96%، ودرجة F1 بنسبة 92.73%.
بالإضافة إلى ذلك، أظهرت دراسة مستخدمين محكومة شملت 100 مشارك أن النموذج يعمل كأداة تعليمية فعالة، حيث حققت المجموعة التجريبية دقة بنسبة 89.4% في الكشف عن المعلومات المضللة مقارنة بـ 74.2% في المجموعة الضابطة. كما أظهر المشاركون في المجموعة التجريبية زيادة في الثقة وتقليل في وقت اتخاذ القرار. تسلط النتائج الضوء على إمكانيات النموذج ليس فقط من حيث الفعالية التقنية في الكشف عن المعلومات المضللة ولكن أيضًا في تعزيز مهارات التفكير النقدي الضرورية لمحو الأمية الرقمية، مما يبرز الدور التحويلي لتقنيات الذكاء الاصطناعي المتقدمة في مكافحة المعلومات المضللة في العصر الرقمي.
الطرق
في هذه الدراسة، استخدم المؤلفون نموذج BERT الأساسي، الذي يتميز بـ 12 طبقة ترميز، و768 وحدة مخفية، وحوالي 110 مليون معلمة، لتحليل مجموعة بيانات تتكون من 20,176 عنصر أخبار مزيفة و19,764 عنصر أخبار حقيقية مأخوذة من منصات التواصل الاجتماعي مثل تويتر. استخدم الإعداد التجريبي أجهزة عالية الأداء، بما في ذلك معالج Intel Xeon وNVIDIA Tesla V100 GPU، بالإضافة إلى 256 جيجابايت من ذاكرة الوصول العشوائي و2 تيرابايت من تخزين NVMe SSD، مما يسهل التدريب والتقييم الفعال للنموذج على نظام تشغيل Ubuntu 20.04 LTS مع TensorFlow 2.4.0. تم تدريب النموذج بحجم دفعة قدره 32، ومعدل تعلم قدره $2 \times 10^{-5}$، ومُحسّن AdamW على مدى 5 عصور، مع تضمين معدل تسرب قدره 0.1 لتخفيف الإفراط في التكيف. تم تقسيم مجموعة البيانات إلى مجموعات تدريب، والتحقق، والاختبار باستخدام أخذ عينات هرمية للحفاظ على توزيع الفئات، وشمل المعالجة المسبقة دمج عناوين الأخبار مع محتواها وتوليد ميزات مثل قطبية المشاعر وعدد الكلمات.
لتقييم تأثير نموذج BERT-LSTM على سلوك المستخدم، تم إجراء دراسة مستخدمين محكومة مع 100 مشارك، مقسمة إلى مجموعة تجريبية (N = 50) تستخدم النموذج ومجموعة ضابطة (N = 50) بدونها. قام المشاركون بتقييم 100 منشور على وسائل التواصل الاجتماعي، مقسمة بالتساوي بين المعلومات الصحيحة والخاطئة. تضمنت المقاييس الرئيسية التي تم قياسها دقة الكشف، ووقت اتخاذ القرار، والثقة في الحكم، بالإضافة إلى تحليل معدلات الخطأ ومعدلات الإيجابيات/السلبيات الكاذبة لتوفير تقييم شامل لفعالية النموذج في مساعدة تقييمات المستخدمين لصدق المعلومات.
النتائج
تسلط نتائج هذه الدراسة الضوء على فعالية نموذج BERT + LSTM في الكشف عن المعلومات المضللة على الشبكات الاجتماعية المحمولة، حيث حقق دقة عالية، واسترجاع، ودرجات F1. يستفيد هذا النموذج من الفهم الدلالي العميق لـ BERT جنبًا إلى جنب مع قدرات معالجة البيانات التسلسلية لـ LSTM، مما يحاكي بشكل فعال العمليات المعرفية المستخدمة في تقييم السرد. مثل هذا الجمع بارع بشكل خاص في تحديد المعلومات المضللة التي تستخدم إشارات لغوية دقيقة أو تلاعب بالمعلومات على مر الزمن.
من منظور عملي، تضع كفاءة النموذج في موقع جيد للتطبيقات في الوقت الحقيقي، وهو أمر حاسم نظرًا للتوزيع السريع للمعلومات المضللة على وسائل التواصل الاجتماعي. يمكن أن تؤدي التأخيرات في الكشف إلى تفاقم تأثير المعلومات الخاطئة، مما يجعل التدخل الفوري أمرًا ضروريًا. بالإضافة إلى ذلك، تقترح الدراسة أن دمج نموذج BERT + LSTM في برامج محو الأمية الرقمية يمكن أن يعزز بشكل كبير من قدرات المستخدمين على تقييم المعلومات عبر الإنترنت بشكل نقدي. تشير النتائج التجريبية إلى أن المشاركين في المجموعة التجريبية أظهروا مستويات ثقة أعلى، وفائدة متصورة، وفعالية تعلم مقارنة بالمجموعة الضابطة، مما يبرز إمكانيات النموذج في تعزيز جمهور أكثر إلمامًا.
المناقشة
تؤكد قسم المناقشة في ورقة البحث على التحدي الحاسم الذي تطرحه المعلومات المضللة في الشبكات الاجتماعية المحمولة وضرورة استجابة متعددة الأوجه تدمج بين الأساليب الحسابية المتقدمة ومبادرات محو الأمية الرقمية. تسلط مراجعة الأدبيات الضوء على تطور تقنيات الكشف عن المعلومات المضللة، حيث انتقلت من طرق التعلم الآلي التقليدية إلى أساليب التعلم العميق المتطورة، مع التركيز بشكل خاص على دور نماذج مثل BERT وLSTM. لقد أثبت التدريب ثنائي الاتجاه وفهم السياق لنموذج BERT فعاليته في تطبيقات متعددة، بما في ذلك الكشف عن المعلومات المضللة، محققًا معدلات دقة عالية. يهدف نموذج BERT-LSTM الهجين المقترح إلى الاستفادة من نقاط القوة في كلا الهيكلين لتعزيز الكشف عن المعلومات المضللة من خلال التقاط الفروق السياقية والأنماط التسلسلية، والتي غالبًا ما يتم تجاهلها من قبل النماذج التقليدية.
بالإضافة إلى ذلك، تؤكد المناقشة على أهمية محو الأمية الرقمية في مكافحة المعلومات المضللة، مشيرة إلى أن التدخلات التعليمية تحسن بشكل كبير من قدرة الأفراد على تمييز المعلومات الموثوقة. يمثل دمج نماذج الكشف المدعومة بالذكاء الاصطناعي مع برامج محو الأمية الرقمية طريقًا واعدًا لتعزيز مهارات التفكير النقدي بين المستخدمين. تهدف الأسئلة البحثية المطروحة إلى استكشاف فعالية نموذج BERT-LSTM في الكشف عن المعلومات المضللة مقارنة بالطرق التقليدية، وقدرته على الكشف المبكر دون بيانات تفاعل المستخدم، وإمكاناته كأداة تعليمية لتحسين محو الأمية الرقمية. بشكل عام، تسعى الدراسة إلى سد الفجوة بين التقدم التكنولوجي في الكشف عن المعلومات المضللة والاستراتيجيات العملية لتعزيز معرفة المستخدمين في العصر الرقمي.
القيود
يقدم نموذج BERT + LSTM تقدمًا ملحوظًا في الكشف عن المعلومات المضللة؛ ومع ذلك، فإنه مقيد بالاعتماد على بيانات أحادية النمط، بشكل أساسي النص. هذه القيود مهمة حيث غالبًا ما تظهر المعلومات المضللة في تنسيقات متعددة الأنماط، مثل الصور، ومقاطع الفيديو، والصوت، خاصة على منصات التواصل الاجتماعي. وبالتالي، قد لا يعالج النموذج بشكل كافٍ تعقيدات المعلومات المضللة في أشكالها المختلفة. بينما يوازن بين الأداء والكفاءة، قد يتفوق عليه نماذج اللغة الكبيرة الأكثر تطورًا (LLMs) مثل GPT-3 أو T5 في سياقات معينة، على الرغم من أن هذه البدائل أقل قابلية للتطبيق للتطبيقات في الوقت الحقيقي بسبب متطلباتها الحسابية العالية.
يجب أن تركز الأبحاث المستقبلية على دمج البيانات متعددة الأنماط في الكشف عن المعلومات المضللة، ربما من خلال نماذج هجينة تجمع بين تحليل النصوص ومعالجة الصور ومقاطع الفيديو. بالإضافة إلى ذلك، يمكن أن تعزز تقنيات التعلم الانتقالي والتعديل من قدرة النموذج على التكيف عبر أنواع وسائط متنوعة. قد توفر دراسة الشبكات العصبية القائمة على الرسوم البيانية (GNNs) أيضًا رؤى حول العلاقات بين الكيانات على وسائل التواصل الاجتماعي، مما يحسن من فعالية النموذج في فهم ديناميات المعلومات المضللة. أخيرًا، يمكن أن يؤدي توسيع نموذج BERT + LSTM لاستيعاب السياقات اللغوية والثقافية المتقاطعة إلى تعزيز قابليته العالمية، مع معالجة القضية المستمرة للمعلومات المضللة عبر لغات وثقافات مختلفة. بشكل عام، يعد التحسين المستمر وتوسيع هذا النموذج أمرًا ضروريًا لتطوير أدوات فعالة لمكافحة المعلومات المضللة وتعزيز محو الأمية الرقمية في عالم متصل.
DOI: https://doi.org/10.1038/s41598-025-85308-4
PMID: https://pubmed.ncbi.nlm.nih.gov/39774143
Publication Date: 2025-01-07
Author(s): Jun Wang et al.
Primary Topic: Misinformation and Its Impacts
Overview
The study addresses the urgent issue of misinformation proliferation in mobile social networks, which poses risks to social stability, public health, and democratic processes. It introduces a hybrid model that integrates Bidirectional Encoder Representations from Transformers (BERT) with Long Short-Term Memory (LSTM) networks, aimed at enhancing the early detection of misinformation based solely on textual content. The BERT-LSTM model demonstrated impressive performance metrics, achieving an accuracy of 93.51%, a recall of 91.96%, and an F1 score of 92.73%.
Additionally, a controlled user study involving 100 participants indicated that the model serves as an effective educational tool, with the experimental group achieving 89.4% accuracy in misinformation detection compared to 74.2% in the control group. Participants in the experimental group also exhibited increased confidence and reduced decision-making time. The findings highlight the model’s potential not only for technical efficacy in misinformation detection but also for fostering critical thinking skills essential for digital literacy, thereby underscoring the transformative role of advanced AI techniques in combating misinformation in the digital age.
Methods
In this study, the authors employed the BERT base model, characterized by 12 encoder layers, 768 hidden units, and approximately 110 million parameters, to analyze a dataset comprising 20,176 fake news items and 19,764 real news items sourced from social media platforms like Twitter. The experimental setup utilized high-performance hardware, including an Intel Xeon CPU and an NVIDIA Tesla V100 GPU, alongside 256 GB of RAM and 2 TB of NVMe SSD storage, facilitating efficient training and evaluation of the model on an Ubuntu 20.04 LTS operating system with TensorFlow 2.4.0. The model was trained with a batch size of 32, a learning rate of $2 \times 10^{-5}$, and the AdamW optimizer over 5 epochs, incorporating a dropout rate of 0.1 to mitigate overfitting. The dataset was partitioned into training, validation, and test sets using hierarchical sampling to maintain class distribution, and preprocessing included concatenating news titles with their content and generating features such as sentiment polarity and word count.
To assess the impact of the BERT-LSTM model on user behavior, a controlled user study was conducted with 100 participants, split into an experimental group (N = 50) using the model and a control group (N = 50) without it. Participants evaluated 100 social media posts, evenly divided between true and false information. Key metrics measured included detection accuracy, decision-making time, and confidence in judgment, alongside an analysis of error rates and false positive/negative rates to provide a comprehensive evaluation of the model’s effectiveness in aiding users’ assessments of information veracity.
Results
The results of this study highlight the effectiveness of the BERT + LSTM model in detecting misinformation on mobile social networks, achieving high accuracy, recall, and F1 scores. This model leverages BERT’s deep semantic understanding alongside LSTM’s sequential data processing capabilities, effectively mimicking cognitive processes used in narrative evaluation. Such a combination is particularly adept at identifying misinformation that employs subtle linguistic cues or information manipulation over time.
From a practical standpoint, the model’s efficiency positions it well for real-time applications, which is crucial given the rapid dissemination of misinformation on social media. Delays in detection can exacerbate the impact of false information, making timely intervention essential. Additionally, the study suggests that integrating the BERT + LSTM model into digital literacy programs could significantly enhance users’ abilities to critically assess online information. The experimental results indicate that participants in the experimental group demonstrated higher confidence levels, perceived usefulness, and learning effectiveness compared to the control group, underscoring the model’s potential to foster a more informed public.
Discussion
The discussion section of the research paper emphasizes the critical challenge posed by misinformation in mobile social networks and the necessity for a multifaceted response that integrates advanced computational methods with digital literacy initiatives. The literature review highlights the evolution of misinformation detection techniques, transitioning from traditional machine learning methods to sophisticated deep learning approaches, particularly emphasizing the role of models like BERT and LSTM. The BERT model’s bidirectional training and contextual understanding have proven effective in various applications, including misinformation detection, achieving high accuracy rates. The proposed BERT-LSTM hybrid model aims to leverage the strengths of both architectures to enhance the detection of misinformation by capturing contextual nuances and sequential patterns, which are often overlooked by conventional models.
Additionally, the discussion underscores the importance of digital literacy in combating misinformation, noting that educational interventions significantly improve individuals’ ability to discern credible information. The integration of AI-powered detection models with digital literacy programs represents a promising avenue for enhancing critical thinking skills among users. The research questions posed aim to explore the effectiveness of the BERT-LSTM model in detecting misinformation compared to traditional methods, its capability for early detection without user engagement data, and its potential as an educational tool for improving digital literacy. Overall, the study seeks to bridge the gap between technological advancements in misinformation detection and practical strategies for enhancing user literacy in the digital age.
Limitations
The BERT + LSTM model presents notable advancements in misinformation detection; however, it is constrained by its reliance on single-modality data, primarily text. This limitation is significant as misinformation often manifests in multimodal formats, such as images, videos, and audio, particularly on social media platforms. Consequently, the model may not adequately address the complexities of misinformation in its various forms. While it balances performance and efficiency, it may still be outperformed by more sophisticated large language models (LLMs) like GPT-3 or T5 in specific contexts, although these alternatives are less feasible for real-time applications due to their high computational requirements.
Future research should focus on integrating multimodal data into misinformation detection, potentially through hybrid models that combine text analysis with image and video processing. Additionally, exploring transfer learning and fine-tuning techniques could enhance the model’s adaptability across diverse media types. Investigating graph-based neural networks (GNNs) may also provide insights into the relationships among entities on social media, further improving the model’s efficacy in understanding misinformation dynamics. Lastly, expanding the BERT + LSTM model to accommodate cross-linguistic and cross-cultural contexts could enhance its global applicability, addressing the pervasive issue of misinformation across different languages and cultures. Overall, continued refinement and expansion of this model are essential for developing effective tools to combat misinformation and promote digital literacy in a connected world.
