DOI: https://doi.org/10.1007/s13278-025-01455-6
تاريخ النشر: 2025-04-22
المؤلف: Mayank Kumar Jain وآخرون
الموضوع الرئيسي: المعلومات المضللة وتأثيراتها
نظرة عامة
تقدم ورقة البحث إطارًا جديدًا لاكتشاف الأخبار المزيفة، يسمى HHO-CNN-BiLSTM، والذي يدمج تحسين صقور هاريس (HHO) لاختيار الميزات مع نموذج هجين يجمع بين الشبكات العصبية التلافيفية (CNN) والشبكات الذاكرة طويلة وقصيرة المدى ثنائية الاتجاه (BiLSTM) للتصنيف. تشمل المنهجية أربع مراحل: إعداد البيانات، استخراج الميزات، اختيار الميزات، والتصنيف. يتم استخدام نهج معالجة البيانات على مرحلتين لتعزيز جودة المدخلات، ويتضمن استخراج الميزات اللغوية وتحسين تضمين الكلمات. تنجح الدراسة في استخراج ثمانين ميزة لغوية، وتحسينها باستخدام HHO، وتحقق دقة تصنيف مثيرة للإعجاب تبلغ 98.89%، 98.25%، 98.56%، و90.26% على مجموعات بيانات ISOT وKaggle وConFake وMcIntire، على التوالي، متجاوزة الطرق الحالية الرائدة.
في الختام، يُظهر إطار HHO-CNN-BiLSTM تقدمًا كبيرًا في اكتشاف الأخبار المزيفة من خلال عمليات اختيار الميزات والتصنيف الفعالة. ستركز الأعمال المستقبلية على توسيع قابلية تطبيق النموذج لاكتشاف الأخبار المزيفة المعتمدة على الصور والتحقق من أدائه على مجموعات بيانات نصية أكبر لتقييم قابلية التوسع. بالإضافة إلى ذلك، يُقترح تطوير وظائف لياقة أكثر تعقيدًا لاختيار الميزات لتعزيز دقة التصنيف والموثوقية، مما يوسع من فائدة النموذج عبر سيناريوهات مختلفة لاكتشاف الأخبار المزيفة.
مقدمة
تسلط مقدمة ورقة البحث الضوء على التأثير المزدوج لوسائل الإعلام الرقمية على المجتمع، مشددة على دورها في تعزيز التواصل والوصول إلى المعلومات بينما تسهل أيضًا انتشار الأخبار المزيفة. توضح الحوادث البارزة، مثل جدل Pizzagate وإعدامات الحشود في الهند، العواقب الوخيمة للمعلومات المضللة، التي يمكن أن تؤدي إلى سلوكيات صحية ضارة، وزيادة تردد اللقاحات، وخسائر اقتصادية كبيرة. تصنف الورقة أشكالًا مختلفة من الأخبار المزيفة، بما في ذلك العناوين المضللة، والقصص المفبركة، والمحتوى المُعدل، ونظريات المؤامرة، مما يبرز ضرورة التفكير النقدي واستراتيجيات تقييم المعلومات الفعالة لمكافحة المعلومات المضللة.
لمعالجة التحديات التي تطرحها الأخبار المزيفة، تقترح الدراسة إطارًا جديدًا يدمج خوارزمية تحسين صقور هاريس (HHO) مع بنية الشبكة العصبية التلافيفية-ذاكرة طويلة وقصيرة المدى ثنائية الاتجاه (CNN-BiLSTM) لتعزيز اكتشاف الأخبار المزيفة. يهدف هذا النهج إلى تحسين اختيار الميزات بشكل ديناميكي، مما يحسن أداء النموذج من خلال تحديد الأنماط ذات الصلة في مجموعات البيانات المعقدة. تثير الدراسة أسئلة بحثية حاسمة بشأن فعالية عملية اختيار الميزات باستخدام HHO وأداء النموذج المقترح مقارنةً بالأنظمة الحالية. من خلال تقييم الإطار على مجموعات بيانات متعددة، تهدف الدراسة إلى الإسهام بشكل كبير في مجال اكتشاف الأخبار المزيفة، مما يعزز كل من الدقة وقابلية التفسير مع معالجة قيود المنهجيات الحالية.
طرق
تم هيكلة المنهجية المقترحة لاكتشاف الأخبار المزيفة في هذه الدراسة كمشكلة تصنيف ثنائية، تهدف إلى تصنيف المقالات الإخبارية إلى فئتين: الأخبار الحقيقية (C₁) والأخبار المزيفة (C₂). الهدف هو تطوير دالة رسم بياني \( f: N \rightarrow y \) تقلل من أخطاء التصنيف مع ضمان المتانة وقابلية التوسع، كما أبرز يانغ وآخرون (2018). تتكون المنهجية من أربع مراحل: إعداد البيانات، استخراج الميزات، اختيار الميزات، والتصنيف. تعالج معالجة البيانات الغموض في المحتوى النصي، بينما يتضمن استخراج الميزات تحديد الميزات اللغوية (LFs). يتم استخدام خوارزمية HHO في مرحلة اختيار الميزات للاحتفاظ بأكثر الميزات ذات الصلة، والتي يتم دمجها بعد ذلك مع تضمينات الكلمات وإدخالها في نموذج هجين CNN-BiLSTM للتصنيف. يتم التحقق من فعالية هذا النهج عبر أربع مجموعات بيانات: ISOT وKaggle وConFake وMcIntire.
تستخدم الإعدادات التجريبية Python 3.12، مع تضمين مكتبات مثل Pandas وSciPy وSciKit-Learn وNumPy وMatplotlib وTextblob وTextstat وNLTK. يتم تقديم مواصفات مفصلة لتكوينات الأجهزة والبرامج في الجدول 5، مما يضمن الشفافية وقابلية إعادة إنتاج نتائج البحث.
نتائج
يوفر قسم النتائج تقييمًا مفصلًا للطريقة المقترحة لاكتشاف الأخبار المزيفة، باستخدام مقاييس أداء متنوعة مثل الدقة، والدقة، ودرجة F1، والاسترجاع عبر أربع مجموعات بيانات: ConFake وMcIntire وISOT وKaggle. تُظهر التحليلات فعالية الطريقة في تمييز الأخبار المزيفة، مما يبرز نجاح خوارزمية تحسين صقور هاريس (HHO) في اختيار أكثر الميزات اللغوية (LFs) ذات الصلة المرتبطة باكتشاف الأخبار المزيفة.
تشير النتائج الرئيسية إلى أن خوارزمية HHO تحدد باستمرار الميزات المتعلقة بعدم اليقين، والذاتية، والعاطفة، والتعقيد، والتي تتكرر في المقالات المضللة. تشمل المؤشرات البارزة الاستخدام المتكرر للغة غير المؤكدة (مثل “ربما”، “قد يكون”)، وعلامات الترقيم المبالغ فيها، ومفردات مدفوعة بالعاطفة محددة (مثل الكلمات السلبية مثل “خوف” و”غضب”). تعزز خوارزمية HHO أيضًا أداء النموذج من خلال تقليل التكرار، مما يضمن الاحتفاظ فقط بأكثر الميزات إفادة، مما يحسن دقة التصنيف ويقلل من المتطلبات الحاسوبية. يساهم دمج ميزات إضافية مثل تكرار الكلمات التوقف، واستخدام الضمائر، وعدد الكلمات الفريدة في تعزيز قابلية تفسير النموذج، حيث تُستخدم هذه العناصر غالبًا في السرد المقنع والدرامي النموذجي للأخبار المزيفة. بشكل عام، تسهم مساحة الميزات المنقحة وآلية البحث التكيفية لـ HHO في إطار تصنيف قوي وقابل للتعميم قادر على معالجة الطبيعة المتطورة للأخبار المزيفة عبر مجموعات بيانات متنوعة.
نقاش
يسلط قسم النقاش في ورقة البحث الضوء على الحاجة الملحة لأساليب فعالة لاكتشاف الأخبار المزيفة، مشددًا على الأساليب المتنوعة التي تم استكشافها في الأدبيات. استخدمت دراسات مختلفة الميزات اللغوية (LFs) وتقنيات التعلم الآلي (ML)، محققة درجات متفاوتة من النجاح. على سبيل المثال، حقق ريدي وآخرون (2020) دقة ملحوظة تبلغ 95.49% باستخدام مزيج من LFs والمصنفات، على الرغم من أن دراستهم واجهت قيودًا بسبب مجموعة بيانات صغيرة ومشكلات الإفراط في التكيف. كما أفاد باحثون آخرون، مثل تشو وآخرون (2020) وغرافانيس وآخرون (2019)، بمعدلات دقة عالية ولكن واجهوا تحديات تتعلق بحجم مجموعة البيانات واختيار الميزات. تشير الورقة إلى اتجاه نحو دمج نماذج التعلم العميق (DL) والخوارزميات الميتاهيرستية لتحسين اختيار الميزات وتعزيز أداء التصنيف، مع إظهار التقدم الأخير دقة محسنة من خلال تقنيات مبتكرة مثل خوارزمية تحسين صقور هاريس (HHO).
يحدد المؤلفون قيودًا كبيرة في الدراسات الحالية، بما في ذلك الاعتماد على مجموعات بيانات فردية ونقص تنوع الميزات، مما يعيق التحليل القوي وأداء النموذج. لمعالجة هذه الفجوات، تقترح الدراسة الحالية عملية شاملة لاستخراج الميزات تتضمن 80 LFs، تليها مرحلة تحسين باستخدام HHO لاختيار أكثر الميزات ذات الصلة لنموذج تصنيف CNN-BiLSTM. يهدف هذا النهج إلى تحسين اكتشاف الأخبار المزيفة من خلال الاستفادة من مجموعة أوسع من الميزات اللغوية والهيكلية، مما يسهم في التقدم المستمر للمنهجيات في هذا المجال البحثي الحاسم. تؤكد النتائج على أهمية مجموعات البيانات المتنوعة واختيار الميزات الفعال في تعزيز متانة ودقة أنظمة اكتشاف الأخبار المزيفة.
DOI: https://doi.org/10.1007/s13278-025-01455-6
Publication Date: 2025-04-22
Author(s): Mayank Kumar Jain et al.
Primary Topic: Misinformation and Its Impacts
Overview
The research paper presents a novel framework for fake news detection, termed HHO-CNN-BiLSTM, which integrates Harris Hawks Optimization (HHO) for feature selection with a hybrid model combining Convolutional Neural Networks (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) networks for classification. The methodology encompasses four phases: data preparation, feature extraction, feature selection, and classification. A two-phase data preprocessing approach is employed to enhance input quality, involving linguistic feature extraction and word embedding refinement. The study successfully extracts eighty linguistic features, optimizing them with HHO, and achieves impressive classification accuracies of 98.89%, 98.25%, 98.56%, and 90.26% on the ISOT, Kaggle, ConFake, and McIntire datasets, respectively, surpassing existing state-of-the-art methods.
In conclusion, the HHO-CNN-BiLSTM framework demonstrates significant advancements in fake news detection through its efficient feature selection and classification processes. Future work will focus on extending the model’s applicability to image-based fake news detection and validating its performance on larger textual datasets to evaluate scalability. Additionally, the development of more sophisticated fitness functions for feature selection is proposed to further enhance classification accuracy and robustness, thereby broadening the model’s utility across various fake news detection scenarios.
Introduction
The introduction of the research paper highlights the dual impact of digital media on society, emphasizing its role in enhancing communication and access to information while also facilitating the spread of fake news. Notable incidents, such as the Pizzagate controversy and mob lynchings in India, illustrate the severe consequences of misinformation, which can lead to harmful health behaviors, increased vaccine hesitancy, and significant economic losses. The paper categorizes various forms of fake news, including misleading headlines, fabricated stories, manipulated content, and conspiracy theories, underscoring the necessity for critical thinking and effective information evaluation strategies to combat misinformation.
To address the challenges posed by fake news, the study proposes a novel framework that integrates the Harris Hawks Optimization (HHO) algorithm with a hybrid Convolutional Neural Network-Bidirectional Long Short-Term Memory (CNN-BiLSTM) architecture for enhanced fake news detection. This approach aims to optimize feature selection dynamically, improving model performance by identifying relevant patterns in complex datasets. The study raises critical research questions regarding the effectiveness of the HHO feature selection process and the comparative performance of the proposed model against existing systems. By evaluating the framework on multiple datasets, the research aims to contribute significantly to the field of fake news detection, enhancing both accuracy and interpretability while addressing limitations of current methodologies.
Methods
The proposed methodology for fake news detection in this study is structured as a binary classification problem, aiming to categorize news articles into two classes: real news (C₁) and fake news (C₂). The objective is to develop a mapping function \( f: N \rightarrow y \) that minimizes classification errors while ensuring robustness and scalability, as highlighted by Yang et al. (2018). The methodology consists of four phases: data preparation, feature extraction, feature selection, and classification. Data preprocessing addresses ambiguities in the textual content, while feature extraction involves the identification of linguistic features (LFs). The HHO algorithm is employed in the feature selection phase to retain the most relevant features, which are then integrated with word embeddings and input into a CNN-BiLSTM hybrid model for classification. The effectiveness of this approach is validated across four datasets: ISOT, Kaggle, ConFake, and McIntire.
The experimental setup utilizes Python 3.12, incorporating libraries such as Pandas, SciPy, SciKit-Learn, NumPy, Matplotlib, Textblob, Textstat, and NLTK. Detailed specifications of the hardware and software configurations are provided in Table 5, ensuring transparency and reproducibility of the research findings.
Results
The results section provides a detailed evaluation of the proposed method for detecting fake news, utilizing various performance metrics such as accuracy, precision, F1 score, and recall across four datasets: ConFake, McIntire, ISOT, and Kaggle. The analysis demonstrates the effectiveness of the method in distinguishing fake news, highlighting the success of the Harris Hawks Optimization (HHO) algorithm in selecting the most relevant linguistic features (LFs) associated with fake news detection.
Key findings indicate that the HHO algorithm consistently identifies features related to uncertainty, subjectivity, sentiment, and complexity, which are prevalent in misleading articles. Notable indicators include the frequent use of tentative language (e.g., “maybe,” “perhaps”), exaggerated punctuation, and specific sentiment-driven vocabulary (e.g., negative words like “fear” and “anger”). The HHO algorithm also enhances model performance by reducing redundancy, ensuring that only the most informative features are retained, which improves classification accuracy and decreases computational demands. The incorporation of additional features such as stopword frequency, pronoun usage, and unique word count further enhances the model’s interpretability, as these elements are often employed in persuasive and dramatic narratives typical of fake news. Overall, the refined feature space and adaptive searching mechanism of HHO contribute to a robust and generalizable classification framework capable of addressing the evolving nature of fake news across diverse datasets.
Discussion
The discussion section of the research paper highlights the critical need for effective fake news detection methodologies, emphasizing the diverse approaches explored in the literature. Various studies have employed linguistic features (LFs) and machine learning (ML) techniques, achieving varying degrees of success. For instance, Reddy et al. (2020) achieved a notable accuracy of 95.49% using a combination of LFs and classifiers, although their study faced limitations due to a small dataset and overfitting issues. Other researchers, such as Zhou et al. (2020) and Gravanis et al. (2019), also reported high accuracy rates but encountered challenges related to dataset size and feature selection. The paper notes a trend towards integrating deep learning (DL) models and metaheuristic algorithms to optimize feature selection and enhance classification performance, with recent advancements demonstrating improved accuracy through innovative techniques like the Harris Hawk Optimization (HHO) algorithm.
The authors identify significant limitations in existing studies, including reliance on single datasets and insufficient feature diversity, which hinder robust analysis and model performance. To address these gaps, the current study proposes a comprehensive feature extraction process that includes 80 LFs, followed by an optimization phase using HHO to select the most relevant features for a CNN-BiLSTM classification model. This approach aims to improve the detection of fake news by leveraging a broader set of linguistic and structural features, ultimately contributing to the ongoing advancement of methodologies in this critical area of research. The findings underscore the importance of diverse datasets and effective feature selection in enhancing the robustness and accuracy of fake news detection systems.
