تحليل المشاعر لمراجعات مدمجة بالرموز التعبيرية باستخدام التعلم الآلي وBert
Sentiment analysis of emoji fused reviews using machine learning and Bert

شارك:
المجلة: Scientific Reports، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-92286-0
PMID: https://pubmed.ncbi.nlm.nih.gov/40038511
تاريخ النشر: 2025-03-04
المؤلف: Amit Khan وآخرون
الموضوع الرئيسي: تحليل المشاعر واستخراج الآراء

نظرة عامة

تسلط الأبحاث الضوء على الدور الهام لمعالجة اللغة الطبيعية (NLP) المدعومة بالذكاء الاصطناعي في تحليل ملاحظات العملاء ضمن قطاع الطيران. تؤكد على أن الرموز التعبيرية والوجوه التعبيرية ضرورية لنقل المشاعر الدقيقة وغالبًا ما توفر معلومات أكثر وضوحًا من النص وحده. على الرغم من أهميتها، فإن الممارسات الحالية غالبًا ما تتضمن إزالة هذه العناصر أثناء معالجة البيانات، مما يمكن أن يعيق دقة تحليل المشاعر.

لمعالجة هذه المشكلة، يقترح المؤلفون نهجًا جديدًا يدمج الرموز التعبيرية والوجوه التعبيرية في تحليل المشاعر لمجموعة بيانات تغريدات شركات الطيران الأمريكية. باستخدام مجموعة متنوعة من مصنفات التعلم الآلي ونموذج BERT، حقق أسلوبهم دقة مثيرة للإعجاب بلغت 92%، متجاوزًا الأطر الحالية الرائدة بنسبة 9%. تؤكد هذه الأبحاث على ضرورة دمج الرموز التعبيرية في تحليل المشاعر لتعزيز فهم آراء وتجارب العملاء.

طرق

في هذا القسم، يوضح المؤلفون الطرق المستخدمة لاستخراج الميزات والتقييم التجريبي لتحليل المشاعر على مراجعات شركات الطيران. استخدموا تكرار المصطلحات – تكرار الوثائق العكسي (TF-IDF) وWord2Vec لمختلف مصنفات التعلم الآلي (ML)، بينما استغل نموذج BERT تقنيات التضمين الخاصة به. يتم تعريف صيغة TF-IDF على أنها \( \text{TF-IDF}(t, d) = \text{TF}(t, d) \cdot \log\left(\frac{N}{\text{DF}(t) + 1}\right) \)، حيث يمثل TF تكرار المصطلح وDF يدل على تكرار الوثيقة. تتكون مجموعة البيانات من 14,460 مراجعة مصنفة إلى مشاعر إيجابية ومحايدة وسلبية، مع وجود 6% ملحوظة تحتوي على رموز تعبيرية. استخدم المؤلفون مصنفات تشمل بايز المتعدد الحدود (MNB)، وغابة عشوائية (RF)، وآلة الدعم الناقل (SVM)، والانحدار اللوجستي (LR)، جنبًا إلى جنب مع نموذج BERT-base، مع تقسيمات التدريب/الاختبار بنسبة 80:20 لـ ML و80:10:10 لـ BERT.

أشارت النتائج التجريبية إلى أن مصنف RF حقق أعلى دقة بلغت 85% في كل من المتغيرين II وIII عند اعتبار الرموز التعبيرية وزيادة البيانات، بينما تبعته MNB بدقة بلغت 80% و81%. تفوق نموذج BERT على مصنفات ML، محققًا دقة بلغت 91% و92% في المتغيرين II وIII، على التوالي. قارن المؤلفون نتائجهم مع الأدبيات الحالية، موضحين أن النماذج المقترحة تفوقت على الأساليب الرائدة بنسبة تقارب 2% مع RF وحوالي 9% مع BERT في المتغير III. يقترحون أن العمل المستقبلي يمكن أن يوسع هذا النهج ليشمل مجموعات بيانات متنوعة لتطبيق أوسع في تحليل المشاعر عبر المجالات.

نقاش

تستعرض قسم النقاش في الورقة الأساليب الحالية لتحليل المشاعر (SA) في قطاع الطيران، مع التركيز على دمج مراجعات العملاء وبيانات وسائل التواصل الاجتماعي، مع التركيز بشكل خاص على دور الرموز التعبيرية والوجوه التعبيرية. بينما استبعدت الدراسات السابقة غالبًا هذه العناصر، فإن هذا البحث رائد في تضمين الرموز التعبيرية، التي أظهرت أنها تعزز تصنيف المشاعر. يشير المؤلفون إلى منهجيات مختلفة، مثل قاموس مشاعر الرموز التعبيرية الذي طوره ألفريهات وآخرون ونموذج CEmo-LSTM بواسطة ليو وآخرون، والذي يدمج بفعالية ميزات الرموز التعبيرية في أطر تحليل المشاعر. على الرغم من التقدم، لا تزال العديد من النماذج التقليدية تعالج الرموز التعبيرية كضوضاء، وهو ما تهدف هذه الدراسة إلى تصحيحه من خلال الاستفادة منها لتحسين رؤى رضا العملاء.

يستخدم النهج المقترح مجموعة بيانات شاملة من قطاع الطيران، مستفيدًا من نمو وسائل التواصل الاجتماعي للحصول على الملاحظات. يقدم المؤلفون نموذج تحويل يستبدل الرموز التعبيرية بتمثيلات نصية باستخدام كل من قاموس مدمج وقاموس مخصص يسمى “EMOJIXT.” يعالج هذا الأسلوب التحديات المتعلقة بعدم توازن البيانات من خلال تطبيق تقنية زيادة العينة للأقليات الاصطناعية (SMOTE). تستخدم الدراسة مصنفات تعلم آلي متنوعة، بما في ذلك الغابة العشوائية وبايز المتعدد الحدود، جنبًا إلى جنب مع نموذج BERT المعدل لتحسين دقة اكتشاف المشاعر. تشير النتائج إلى أن دمج الرموز التعبيرية يحسن بشكل كبير أداء نماذج تصنيف المشاعر، مما يظهر إمكانية فهم أكثر دقة لمشاعر العملاء في صناعة الطيران.

القيود

تسلط هذه القسم الضوء على قيود تقنيات استخراج الميزات المختلفة لتحليل المشاعر (SA). يقدم قاموس مشاعر الرموز التعبيرية (Emo-SL) مع طرق التعلم الآلي مثل آلات الدعم الناقل (SVM)، وأقرب الجيران (KNN)، والغابة العشوائية (RF)، وبايز البسيط (NB) مزايا في تصنيف المشاعر من خلال دمج ميزات الرموز التعبيرية جنبًا إلى جنب مع العناصر اللغوية التقليدية مثل تصنيف أجزاء الكلام (POS) وعدد الكلمات. ومع ذلك، فإن هذا النهج مقيد باعتماده على المشاعر المستندة إلى الرموز التعبيرية، والتي تعتمد بطبيعتها على اللغة وقد لا تعالج بشكل كافٍ الاختلافات اللهجية أو غير الرسمية في اللغة.

بالإضافة إلى ذلك، فإن استخدام الشبكات العصبية طويلة وقصيرة الأجل (LSTM) مع آلية انتباه النمط لدمج ميزات النص والرموز التعبيرية يظهر وعدًا في تعزيز تحليل المشاعر. ومع ذلك، يفتقر هذا النظام حاليًا إلى دمج تضمينات الكلمات المتقدمة مثل Word2Vec وFastText أو BERT، والتي يمكن أن تحسن بشكل كبير عملية استخراج الميزات. أخيرًا، يواجه تحليل المشاعر المستند إلى الرموز التعبيرية باستخدام نهج قاموس SentiWordNet أيضًا قيودًا مماثلة في التقاط الطيف الكامل من الفروق الدقيقة في المشاعر الموجودة في سياقات لغوية متنوعة.

Journal: Scientific Reports, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-92286-0
PMID: https://pubmed.ncbi.nlm.nih.gov/40038511
Publication Date: 2025-03-04
Author(s): Amit Khan et al.
Primary Topic: Sentiment Analysis and Opinion Mining

Overview

The research highlights the significant role of Natural Language Processing (NLP) powered by Artificial Intelligence in analyzing customer feedback within the aviation sector. It emphasizes that emojis and emoticons are crucial for conveying nuanced sentiments and often provide more explicit information than text alone. Despite their importance, current practices frequently involve the removal of these elements during data preprocessing, which can hinder sentiment analysis accuracy.

To address this issue, the authors propose a novel approach that incorporates emojis and emoticons in sentiment analysis of the US Airline tweet dataset. Utilizing various Machine Learning classifiers and the BERT model, their method achieved an impressive accuracy of 92%, surpassing existing state-of-the-art frameworks by 9%. This research underscores the necessity of integrating emotive symbols in sentiment analysis to enhance the understanding of customer opinions and experiences.

Methods

In this section, the authors detail the methods employed for feature extraction and experimental evaluation of sentiment analysis on airline reviews. They utilized Term Frequency-Inverse Document Frequency (TF-IDF) and Word2Vec for various machine learning (ML) classifiers, while the BERT model leveraged its own embedding techniques. The TF-IDF formula is defined as \( \text{TF-IDF}(t, d) = \text{TF}(t, d) \cdot \log\left(\frac{N}{\text{DF}(t) + 1}\right) \), where TF represents term frequency and DF denotes document frequency. The dataset comprised 14,460 reviews categorized into positive, neutral, and negative sentiments, with a notable 6% containing emojis. The authors employed classifiers including Multinomial Naive Bayes (MNB), Random Forest (RF), Support Vector Machine (SVM), and Logistic Regression (LR), alongside the BERT-base transformer model, with training/testing splits of 80:20 for ML and 80:10:10 for BERT.

The experimental results indicated that the RF classifier achieved the highest accuracy of 85% in both variants II and III when considering emojis and data augmentation, while MNB followed with accuracies of 80% and 81%. The BERT model outperformed the ML classifiers, achieving 91% and 92% accuracy in variants II and III, respectively. The authors compared their results with existing literature, demonstrating that their proposed models outperformed state-of-the-art methods by approximately 2% with RF and around 9% with BERT in variant III. They suggest that future work could expand this approach to include diverse datasets for broader applicability in cross-domain sentiment analysis.

Discussion

The discussion section of the paper reviews existing sentiment analysis (SA) approaches in the airline sector, emphasizing the integration of customer reviews and social media data, particularly focusing on the role of emojis and emoticons. While previous studies have often excluded these elements, this research is pioneering in its inclusion of emojis, which are shown to enhance sentiment classification. The authors reference various methodologies, such as the Emoji Sentiment Lexicon developed by Alfreihat et al. and the CEmo-LSTM model by Liu et al., which effectively incorporate emoji features into sentiment analysis frameworks. Despite the advancements, many traditional models still treat emojis as noise, which this study aims to rectify by leveraging them to improve customer satisfaction insights.

The proposed approach utilizes a comprehensive dataset from the aviation sector, capitalizing on the growth of social media for feedback acquisition. The authors introduce a transformation model that replaces emojis with textual representations using both an inbuilt dictionary and a custom dictionary named “EMOJIXT.” This method addresses challenges related to data imbalance through the application of the Synthetic Minority Oversampling Technique (SMOTE). The study employs various machine learning classifiers, including Random Forest and Multinomial Naive Bayes, alongside a fine-tuned BERT model to enhance sentiment detection accuracy. The results indicate that incorporating emojis significantly improves the performance of sentiment classification models, demonstrating the potential for more nuanced understanding of customer sentiments in the airline industry.

Limitations

The limitations of various feature extraction techniques for sentiment analysis (SA) are highlighted in this section. The Emoji Sentiment Lexicon (Emo-SL) combined with machine learning methods such as Support Vector Machines (SVM), K-nearest neighbors (KNN), Random Forest (RF), and Naive Bayes (NB) offers advantages in sentiment classification by incorporating emoji features alongside traditional linguistic elements like Parts of Speech (POS) tagging and word counts. However, this approach is constrained by its reliance on emoji-based sentiment, which is inherently language-dependent and may not adequately address dialectal or informal variations in language.

Additionally, the use of Long Short-Term Memory (LSTM) networks with a modality attention mechanism for fusing text and emoji features shows promise in enhancing sentiment analysis. Nonetheless, this system currently lacks the integration of advanced word embeddings such as Word2Vec, FastText, or BERT, which could significantly improve the feature extraction process. Lastly, the emoji-based sentiment analysis utilizing the SentiWordNet lexicon approach also faces similar limitations in capturing the full spectrum of sentiment nuances present in diverse linguistic contexts.

شارك: