DOI: https://doi.org/10.1038/s41598-025-22853-y
PMID: https://pubmed.ncbi.nlm.nih.gov/41214037
تاريخ النشر: 2025-11-10
المؤلف: Jing Tan وآخرون
الموضوع الرئيسي: سمات الشخصية وعلم النفس
نظرة عامة
تتناول هذه الدراسة تحدي عدم التوازن في مجموعات بيانات الشخصية، والذي يمكن أن يؤدي إلى أداء متحيز في نماذج التصنيف. يقترح المؤلفون تقنية جديدة تسمى خسارة بؤرية تكيفية مع تقسيم مجموعة بيانات مصنفة حسب الشخصية، تهدف إلى تعزيز التعرف على الشخصية متعددة الأبعاد. من خلال استخدام تقنيات مختلفة للتخفيف من عدم التوازن في الفئات (CIMTs)، بما في ذلك طرق أخذ العينات ودوال الخسارة المتخصصة، تظهر الدراسة أن هذه الأساليب المصنفة تستقر بشكل كبير في الأداء عبر تقسيمات مجموعات البيانات المختلفة. تشير النتائج إلى أن الطريقة المقترحة تحسن الدقة المتوازنة (BA) بنسبة تصل إلى 7% على مجموعة بيانات كاجل و5% على مجموعة بيانات المقالات مقارنة بأساليب التدريب التقليدية، مع تحمل الحد الأدنى من الأعباء الحاسوبية.
في الخاتمة، يؤكد المؤلفون على أهمية تقسيم مجموعة بياناتهم المصنفة حسب الشخصية وتمثيل العلامات، مما يعزز الاستقرار وكفاءة التعلم بنسبة تصل إلى 9%. يبرزون أن خسارتهم البؤرية التكيفية تتفوق على الأساليب التقليدية، محققة تحسناً بنسبة 7% في BA على مجموعة بيانات كاجل و5% على مجموعة بيانات المقالات، متجاوزة حتى النماذج المعتمدة على المحولات. تدعو الدراسة إلى استخدام الدقة المتوازنة كمقياس تقييم رئيسي بدلاً من الدقة العادية ودرجة F1، حيث توفر انعكاسًا أكثر دقة لأداء النموذج في توزيعات الفئات غير المتوازنة. يقترح المؤلفون أن تحسين عمليات التدريب يمكن أن يؤدي إلى فوائد أداء أكبر من مجرد زيادة تعقيد النموذج، وهو ما يكون ذا صلة خاصة بالتطبيقات في البيئات ذات الموارد المحدودة، مثل الروبوتات.
طرق
في هذا القسم، يوضح المؤلفون المنهجية المستخدمة لخوارزمية خسارة بؤرية تكيفية مصممة للتعرف على الشخصية متعددة الأبعاد. تبدأ العملية بإعداد مجموعة البيانات، والتي تشمل تقسيم النصوص، وتنظيف البيانات، واستخراج الميزات باستخدام تقنيات مثل LIWC وTF-IDF. يستخدم المؤلفون نهجًا مصنفًا لتقسيم مجموعة البيانات لمعالجة عدم التوازن في الفئات، مما يضمن أن تكون بيانات التدريب ممثلة لمختلف أنواع الشخصية.
تتضمن عملية التدريب مفتش توازن الدفعات لتقييم توزيع الفئات ومعدل ضبط الوزن الذي يحسب الأوزان الخاصة بكل حالة. يتم دمج هذه الأوزان في دالة خسارة بؤرية تكيفية، والتي تتضمن قياس الصعوبة لتعزيز أداء النموذج. خلال مرحلة الانتشار العكسي، يتم حساب خسارة الانتروبيا المتقاطعة الثنائية (BCE)، ويتم تطبيق عامل بؤري تكيفي لضبط قياس التدرج ديناميكيًا، مما يحسن عملية التعلم لمهمة التعرف على الشخصية متعددة الأبعاد. يتم تمثيل المنهجية بصريًا في الشكل 1، موضحًا خط التدريب الشامل والميزات المبتكرة للخوارزمية المقترحة.
النتائج
في قسم النتائج، يقدم المؤلفون نتائجهم مرتبة في ثلاثة مجالات رئيسية. أولاً، توضح تصنيفات مجموعة البيانات استقرار تقييماتهم عبر بذور عشوائية مختلفة، مما يشير إلى قوة تصميمهم التجريبي. ثانيًا، يقيم أداء CIMT فعالية المنهجيات المقترحة من خلال مقارنتها بتنفيذات مكافئة، مما يوفر رؤى حول أدائها النسبي. أخيرًا، توضح قسم المعايير كيف تقارن أساليب المؤلفين مع النتائج الحالية الرائدة، مما يعرض تقدمهم في هذا المجال. بشكل عام، تؤكد هذه النتائج موثوقية وتنافسية الأساليب المقترحة.
المناقشة
في قسم المناقشة من الورقة، يستكشف المؤلفون الطبيعة متعددة الأوجه للتعبير عن الشخصية من خلال الوسائط البصرية والصوتية والنصية، مؤكدين على عملية الكشف عن الشخصية المعتمدة على النصوص. يبرزون أن الشخصية تتأثر بالعوامل البيولوجية والبيئية، مما يؤدي إلى أنواع شخصية متنوعة. توفر البيانات النصية من وسائل التواصل الاجتماعي مصدرًا غنيًا لتحليل الشخصية، حيث تعكس غالبًا العمليات المعرفية والحالات العاطفية للفرد. يشير المؤلفون إلى نتائج ياركوني حول العلاقة بين استخدام الكلمات وسمات الشخصية، مقترحين أن الآثار الرقمية التي يتركها الأفراد يمكن استغلالها لتطبيقات واعية بالشخصية، خاصة في مجال الروبوتات.
تستعرض الورقة أيضًا نظريات الشخصية المختلفة، بما في ذلك نماذج Big-3 وBig-4 وBig-5 وBig-6، مشيرة إلى العلاقات المحدودة بينها، مما يعقد نقل أطر الشخصية عبر الأنظمة. ينسق المؤلفون عملهم مع نموذج Big-5 من أجل الوضوح. يناقشون التقدم في الخوارزميات الخاصة بالتعرف على الشخصية، بما في ذلك تقنيات استخراج الميزات مثل تطبيع النصوص، وتقسيم النصوص، ودمج الميزات النفسية اللغوية. يستشهد المؤلفون بدراسات مهمة استخدمت نماذج التعلم الآلي، بما في ذلك CNNs وRNNs، لتحسين دقة توقع الشخصية، مع معالجة التحدي المستمر لعدم التوازن في الفئات في مجموعات البيانات. يقترحون تقنيات جديدة، مثل خسارة بؤرية تكيفية وتوزيع مجموعة بيانات مصنفة حسب الشخصية، لتعزيز أداء النموذج والعدالة في مهام التعرف على الشخصية.
DOI: https://doi.org/10.1038/s41598-025-22853-y
PMID: https://pubmed.ncbi.nlm.nih.gov/41214037
Publication Date: 2025-11-10
Author(s): Jing Tan et al.
Primary Topic: Personality Traits and Psychology
Overview
This research addresses the challenge of class imbalance in personality datasets, which can lead to biased performance in classification models. The authors propose a novel technique called Adaptive Focal Loss with Personality-Stratified Dataset Splitting, aimed at enhancing multi-dimensional personality recognition. By employing various class imbalance mitigation techniques (CIMTs), including sampling methods and specialized loss functions, the study demonstrates that these stratified approaches significantly stabilize performance across different dataset splits. The findings indicate that the proposed method improves balanced accuracy (BA) by up to 7% on the Kaggle dataset and 5% on the Essays dataset compared to traditional training methods, while incurring minimal computational overhead.
In the conclusion, the authors emphasize the importance of their personality-stratified dataset splitting and label representation, which enhance stability and learning efficiency by up to 9%. They highlight that their Adaptive Focal Loss outperforms conventional methods, achieving a 7% improvement in BA on the Kaggle dataset and a 5% improvement on the Essays dataset, even surpassing transformer-based models. The study advocates for the use of balanced accuracy as a primary evaluation metric over regular accuracy and F1 score, as it provides a more accurate reflection of model performance in imbalanced class distributions. The authors suggest that optimizing training processes can yield greater performance benefits than merely increasing model complexity, which is particularly relevant for applications in resource-constrained environments, such as robotics.
Methods
In this section, the authors outline the methodology employed for their adaptive focal loss algorithm tailored for multi-dimensional personality recognition. The process begins with dataset preparation, which includes text tokenization, data cleaning, and feature extraction using techniques such as LIWC and TF-IDF. The authors utilize a stratified approach for dataset splitting to address class imbalance, ensuring that the training data is representative of the various personality types.
The training pipeline incorporates a batch balancing inspector to evaluate class distribution and a weight adjuster that computes instance-specific weights. These weights are integrated into the Adaptive Focal Loss function, which incorporates difficulty scaling to enhance model performance. During the backpropagation phase, the binary cross-entropy (BCE) loss is calculated, and an adaptive focal factor is applied to dynamically adjust the gradient scaling, thereby optimizing the learning process for the multi-dimensional personality recognition task. The methodology is visually represented in Figure 1, illustrating the comprehensive training pipeline and the innovative features of the proposed algorithm.
Results
In the Results section, the authors present their findings organized into three key areas. First, the Dataset Stratification demonstrates the stability of their evaluations across various random seeds, indicating the robustness of their experimental design. Second, the CIMT Performance assesses the effectiveness of the proposed methodologies by comparing them against equivalent implementations, providing insights into their relative performance. Lastly, the Benchmarks section illustrates how the authors’ methods measure up against existing state-of-the-art results, showcasing their advancements in the field. Overall, these results underscore the reliability and competitiveness of the proposed approaches.
Discussion
In the discussion section of the paper, the authors explore the multifaceted nature of personality expression through visual, audio, and textual modalities, emphasizing the practicality of text-based personality detection. They highlight that personality is influenced by biological and environmental factors, leading to varied personality types. Textual data from social media provides a rich source for personality analysis, as it often reflects an individual’s cognitive processes and emotional states. The authors reference Yarkoni’s findings on the correlation between word usage and personality traits, suggesting that the digital footprints left by individuals can be leveraged for personality-aware applications, particularly in robotics.
The paper also reviews various personality theories, including the Big-3, Big-4, Big-5, and Big-6 models, noting the limited correlations among them, which complicates the transfer of personality frameworks across systems. The authors align their work with the Big-5 model for clarity. They discuss advancements in algorithms for personality recognition, including feature extraction techniques such as text normalization, tokenization, and psycholinguistic feature integration. The authors cite significant studies that have utilized machine learning models, including CNNs and RNNs, to improve personality prediction accuracy, while also addressing the persistent challenge of class imbalance in datasets. They propose novel techniques, such as adaptive focal loss and personality-stratified dataset distribution, to enhance model performance and fairness in personality recognition tasks.
