تعلم كيفية صياغة مطالبات متنوعة لتصنيف عواطف الصور
Learning to compose diversified prompts for image emotion classification

شارك:
المجلة: Computational Visual Media، المجلد: 10، العدد: 6
DOI: https://doi.org/10.1007/s41095-023-0389-6
تاريخ النشر: 2024-05-06
المؤلف: Sinuo Deng وآخرون
الموضوع الرئيسي: التعرف على الوجه والتعبيرات

نظرة عامة

تقدم ورقة البحث إطار عمل جديد، PT-DPC، لتصنيف عواطف الصور (IEC)، مستفيدة من قدرات الأساليب المعتمدة على اللغة مثل التدريب المسبق المتباين بين اللغة والصورة (CLIP). تحدد الدراسة ثلاثة تحديات رئيسية في IEC: الفجوة الكبيرة في أهداف التدريب بين التدريب المسبق وIEC، واستخدام مطالبات مشتركة غير مثلى، وتطبيق مطالبات ثابتة عبر الحالات. لمعالجة هذه القضايا، يقدم المؤلفون طريقة ضبط المطالبات التي تتماشى مع أهداف التدريب المسبق لـ CLIP وتولد تلقائيًا مطالبات محددة للحالات بناءً على الفئات ومحتوى الصور. تعمل هذه التنويعات في المطالبات على تقليل الأداء غير المثالي بشكل فعال.

تظهر تقييمات PT-DPC على ستة مجموعات بيانات عاطفية مستخدمة على نطاق واسع تفوقه على الأساليب الحديثة، محققًا تحسينًا في الدقة يصل إلى 9.29% على مجموعة بيانات EmotionROI مع الحد الأدنى من المعلمات الإضافية. تشير النتائج إلى أنه بينما تكتسب الأساليب متعددة الوسائط زخمًا، لا تزال العديد من الدراسات الحالية تركز على الوسائط الفردية. يدعو المؤلفون إلى تطبيق إطار عملهم لإلهام المزيد من الأبحاث في المهام ذات الموارد المحدودة وتطوير استراتيجيات مبتكرة متعددة الوسائط للتحديات التقليدية ذات الوسائط الفردية. الشيفرة الخاصة بـ PT-DPC متاحة للجمهور لمزيد من البحث.

مقدمة

تناقش مقدمة ورقة البحث أهمية التعرف على العواطف في الذكاء الاصطناعي، مشددة على دورها في فهم نوايا البشر من خلال منهجيات متنوعة، بما في ذلك التعلم العميق وتحليل العواطف البصرية. تسلط الضوء على تطور طرق تحليل العواطف، مع التركيز بشكل خاص على تصنيف عواطف الصور (IEC)، الذي يهدف إلى استخراج العواطف من الصور. استخدمت الطرق التقليدية مجموعات تسميات ثابتة وهياكل، لكن التقدمات الأخيرة، مثل التدريب المسبق المتباين بين اللغة والصورة (CLIP)، تقدم بديلاً واعدًا من خلال محاذاة الصور مع أوصاف اللغة الطبيعية، مما يعزز قدرات التعميم.

على الرغم من إمكانيات CLIP لمهام IEC، تحدد الورقة ثلاثة تحديات رئيسية: الفجوة بين التدريب المسبق وIEC بسبب الاستخدام الحصري لبيانات الصور أثناء التدريب، والأداء غير المثالي المرتبط بالمطالبات النصية المصممة يدويًا، والقيود المفروضة على المطالبات الثابتة عبر حالات مختلفة. لمعالجة هذه القضايا، يقترح المؤلفون طريقة جديدة تُسمى ضبط المطالبات مع تكوين مطالبات متنوعة (PT-DPC)، والتي تولد مطالبات فريدة لكل صورة من خلال الاستفادة من الرموز الافتراضية الخاصة بالفئة. تظهر النتائج التجريبية أن PT-DPC يتفوق بشكل كبير على الأساليب الحالية، محققًا تحسينًا ملحوظًا في الدقة على عدة معايير لتصنيف العواطف. كما توضح الدراسة التحديات المتبقية، مثل التحيزات في البيانات وتعقيد الفئات العاطفية، مقترحةً طرقًا للبحث المستقبلي لتعزيز دقة تصنيف العواطف.

الطرق

يتم توضيح إطار عمل PT-DPC المقترح من خلال ثلاث وجهات نظر رئيسية: تعريف المشكلة، تكوين المطالبات المتنوعة، ومنهجية التدريب. تستعرض هذه القسم طرق الأساس المختلفة لتصنيف عواطف الصور، مع تسليط الضوء على كل من الأساليب التقليدية المعتمدة على الميزات المصممة يدويًا وتقنيات التعلم العميق المعاصرة. استخدمت الطرق المبكرة مصنفات بسيطة مع ميزات مصممة يدويًا، بينما استغلت التقدمات الأخيرة الشبكات العصبية التلافيفية (CNNs) لاستخراج ميزات عميقة. تشمل النماذج البارزة DeepSentibank، الذي يركز على تصنيف المشاعر البصرية، والشبكة العصبية المتصلة بالنبض (PCNN) التي تستخدم بنية CNN تقدمية. تشمل المساهمات الأخرى المهمة الشبكة المقترنة ذات الإشراف الضعيف (WSCNet) وفقدان الوعي بفئة العواطف (ECWA)، وكلاهما حقق أداءً رائدًا في مجموعة متنوعة من البيانات.

يقارن القسم أيضًا PT-DPC مع الأساليب التقليدية المعتمدة على اللغة، مؤكدًا على كفاءته في استخدام المعلمات. بينما يظهر نموذج SimEmotion الكامل أداءً متفوقًا، إلا أنه يتطلب أكثر من 351.39 مليون معلمة، وهو ما يزيد بشكل كبير عن عدد المعلمات القليل لـ PT-DPC. يتطلب إصدار التدريب MLP من SimEmotion، على الرغم من أنه مفيد قليلاً على مجموعات بيانات معينة، ما يقرب من 30 مرة من معلمات PT-DPC. طريقة Adapter، المستوحاة من معالجة اللغة الطبيعية، لديها عدد معلمات مشابه لـ PT-DPC لكنها تفتقر إلى التدريب المحدد للمهام، مما يؤدي إلى أداء أقل. يبرز هذا المقارنة كفاءة وفعالية إطار عمل PT-DPC في مهام تصنيف عواطف الصور.

المناقشة

في قسم المناقشة من الورقة، يستكشف المؤلفون مجموعة متنوعة من الأعمال والأساليب ذات الصلة في مجال تمثيل العواطف وتصنيف عواطف الصور. يبرزون نموذجين رئيسيين لتمثيل العواطف: الفضاء العاطفي البُعدي (DES)، وبشكل خاص نموذج الفالنس-التحفيز-الهيمنة (VAD)، وحالات العواطف الفئوية (CES). بينما يوفر DES تمثيلًا مستمرًا للعواطف، فإنه غالبًا ما يكون من الصعب على المستخدمين تفسيره. على العكس، فإن نماذج CES، التي تصنف العواطف إلى أنواع منفصلة، أكثر بديهية ولكن قد تبسط التعقيد العاطفي بشكل مفرط. يقترح المؤلفون طريقة جديدة لتصنيف عواطف الصور بناءً على نموذج CES، مؤكدين على أهمية استخراج ميزات تمييزية من الصور لتعزيز أداء التصنيف.

كما يناقش المؤلفون التقدمات في النماذج المدربة مسبقًا على نطاق واسع وطرق ضبط المطالبات في معالجة اللغة الطبيعية (NLP). يشيرون إلى أنه بينما حسنت الشبكات العصبية العميقة، مثل CNNs، الأداء في مهام متنوعة، فإنها غالبًا ما تتطلب بيانات موسومة بشكل كبير ويمكن أن تعاني من الإفراط في التكيف. لمعالجة هذه التحديات، تقدم الورقة إطار عمل PT-DPC، الذي يستخدم نموذجًا مدربًا مسبقًا لتوليد مطالبات متنوعة تتناسب مع محتوى الصورة المحدد وفئات العواطف. لا تعزز هذه الطريقة دقة التصنيف عبر مجموعات بيانات متعددة فحسب، بل تظهر أيضًا فعالية دمج المعلومات المحددة للحالات والمعلومات المحددة للفئة في ضبط المطالبات. بشكل عام، يتفوق الأسلوب المقترح على التقنيات الحديثة الحالية، مما يبرز إمكانيته في تعزيز مهام تصنيف العواطف متعددة الوسائط.

Journal: Computational Visual Media, Volume: 10, Issue: 6
DOI: https://doi.org/10.1007/s41095-023-0389-6
Publication Date: 2024-05-06
Author(s): Sinuo Deng et al.
Primary Topic: Face and Expression Recognition

Overview

The research paper presents a novel framework, PT-DPC, for image emotion classification (IEC), leveraging the capabilities of language-supervised methods like contrastive language-image pretraining (CLIP). The study identifies three primary challenges in IEC: the significant training objective gap between pretraining and IEC, the use of shared suboptimal prompts, and the application of invariant prompts across instances. To address these issues, the authors introduce a prompt-tuning method that aligns with CLIP’s pretraining objectives and automatically generates instance-specific prompts based on the categories and content of the images. This diversification of prompts effectively mitigates suboptimal performance.

The evaluation of PT-DPC on six widely used affective datasets demonstrates its superiority over state-of-the-art methods, achieving up to a 9.29% accuracy improvement on the EmotionROI dataset with minimal additional parameters. The findings suggest that while multimodal approaches are gaining traction, many existing studies still focus on single modalities. The authors advocate for the application of their framework to inspire further research in resource-constrained tasks and the development of innovative multimodal strategies for traditional single-modal challenges. The code for PT-DPC is publicly accessible for further research.

Introduction

The introduction of the research paper discusses the significance of emotion recognition in artificial intelligence, emphasizing its role in understanding human intentions through various methodologies, including deep learning and visual emotion analysis. It highlights the evolution of emotion analysis methods, particularly focusing on image emotion classification (IEC), which aims to extract emotions from images. Traditional approaches have utilized fixed label sets and architectures, but recent advancements, such as contrastive language-image pretraining (CLIP), offer a promising alternative by aligning images with natural language descriptions, thereby enhancing generalization capabilities.

Despite the potential of CLIP for IEC tasks, the paper identifies three major challenges: the disparity between pretraining and IEC due to the exclusive use of image data during training, the suboptimal performance linked to manually designed text prompts, and the limitations of shared invariant prompts across different instances. To address these issues, the authors propose a novel method called prompt tuning with diversified prompt composition (PT-DPC), which generates unique prompts for each image by leveraging class-specific virtual tokens. Experimental results demonstrate that PT-DPC significantly outperforms existing methods, achieving a notable accuracy improvement on multiple emotion classification benchmarks. The study also outlines remaining challenges, such as domain biases in datasets and the complexity of emotional categories, suggesting avenues for future research to enhance emotion classification accuracy.

Methods

The proposed PT-DPC framework is outlined through three key perspectives: problem definition, diversified prompt composition, and training methodology. The section reviews various baseline methods for image-emotion classification, highlighting both traditional handcrafted feature-based approaches and contemporary deep learning techniques. Early methods utilized simple classifiers with hand-crafted features, while recent advancements have leveraged convolutional neural networks (CNNs) to extract deep features. Notable models include DeepSentibank, which focuses on visual sentiment classification, and the pulse-coupled neural network (PCNN) that employs a progressive CNN architecture. Other significant contributions include the weakly supervised coupled network (WSCNet) and emotion class-wise aware loss (ECWA), both achieving state-of-the-art (SOTA) performance on various datasets.

The section further compares PT-DPC with traditional language-supervised methods, emphasizing its efficiency in parameter utilization. While the full-version SimEmotion model demonstrates superior performance, it requires over 351.39 million parameters, significantly more than PT-DPC’s minimal parameter count. The MLP-training version of SimEmotion, although slightly advantageous on specific datasets, still necessitates nearly 30 times the parameters of PT-DPC. The Adapter method, inspired by natural language processing, has a parameter count similar to PT-DPC but lacks task-specific training, resulting in inferior performance. This comparison underscores the efficiency and effectiveness of the PT-DPC framework in image-emotion classification tasks.

Discussion

In the discussion section of the paper, the authors explore various related works and methodologies in the field of emotion representation and image emotion classification. They highlight two primary models for emotion representation: the dimensional emotion space (DES), particularly the valence-arousal-dominance (VAD) model, and categorical emotion states (CES). While DES provides a continuous representation of emotions, it is often challenging for users to interpret. Conversely, CES models, which categorize emotions into discrete types, are more intuitive but may oversimplify emotional complexity. The authors propose a novel image emotion classification method based on the CES model, emphasizing the importance of extracting discriminative features from images to enhance classification performance.

The authors also discuss advancements in large-scale pretrained models and prompt-tuning methods in natural language processing (NLP). They note that while deep neural networks, such as CNNs, have improved performance in various tasks, they often require extensive labeled data and can suffer from overfitting. To address these challenges, the paper introduces the PT-DPC framework, which utilizes a pretrained model to generate diversified prompts that are tailored to specific image content and emotion categories. This approach not only enhances classification accuracy across multiple datasets but also demonstrates the effectiveness of integrating instance-specific and class-specific information in prompt tuning. Overall, the proposed method outperforms existing state-of-the-art techniques, showcasing its potential for advancing multimodal emotion classification tasks.

شارك: