نماذج اللغة الكبيرة يمكن أن تمكن التحليل الموضوعي الاستقرائي لمجموعة بيانات وسائل التواصل الاجتماعي في طلب واحد: دراسة تحقق بشرية
Large Language Models Can Enable Inductive Thematic Analysis of a Social Media Corpus in a Single Prompt: Human Validation Study

شارك:
المجلة: JMIR Infodemiology، المجلد: 4
DOI: https://doi.org/10.2196/59641
PMID: https://pubmed.ncbi.nlm.nih.gov/39207842
تاريخ النشر: 2024-07-01
المؤلف: Michael Deiner وآخرون
الموضوع الرئيسي: المعلومات المضللة وتأثيراتها

نظرة عامة

تبحث الدراسة في جدوى استخدام نماذج اللغة الكبيرة التوليدية (LLMs) لاختيار نماذج الموضوع والتحليل الموضوعي الاستقرائي للمنشورات على وسائل التواصل الاجتماعي المتعلقة بالصحة العامة. نظرًا للتحديات التي تواجه التحليل اليدوي من قبل خبراء مدربين، تقارن الدراسة أداء ثلاثة نماذج LLMs—GPT4-32K وClaudeinstant-100K وClaude-2-100K—مع التحليلات البشرية من دراسة سابقة حول بلاغة اللقاح. تشير النتائج إلى أن جميع نماذج LLMs الثلاثة تلخصت بفعالية المحتوى وحددت الموضوعات ذات الصلة، مع دلالة إحصائية (P<0.001) تشير إلى احتمال أعلى لتضمين مجالات المحتوى الأعلى تقييمًا من قبل البشر في تصنيفاتها. بينما حددت نماذج LLMs موضوعات كانت عمومًا قابلة للمقارنة مع تلك التي اعترف بها الخبراء البشر، لوحظت تباينات بين نماذج LLMs المختلفة وعبر التحليلات المتكررة لنفس النموذج. على الرغم من عدم تكرار عمق التحليل البشري بشكل متسق، اعتبرت الموضوعات التي أنشأتها نماذج LLMs معقولة وذات صلة من قبل خبراء الموضوع. تسلط النتائج الضوء على إمكانيات نماذج LLMs في الاستماع الآلي لوسائل التواصل الاجتماعي في الصحة العامة، على الرغم من أن التحقق الإضافي ضروري لتعزيز دقتها وموثوقيتها في استخراج الموضوعات الدقيقة.

مقدمة

تناقش مقدمة هذه الورقة البحثية إمكانيات منصات وسائل التواصل الاجتماعي لتعزيز فهم الصحة العامة، لا سيما فيما يتعلق بمشاعر اللقاح، وهو أمر حيوي للوقاية من الأمراض ورفاهية الجمهور. على الرغم من وفرة المحتوى الذي ينشئه المستخدمون المتاح، فإن تحليل هذه البيانات غير المنظمة غالبًا ما يكون كثيف العمالة ويستغرق وقتًا طويلاً، مما يدفع الباحثين إلى استخدام طرق مثل معالجة اللغة الطبيعية (NLP) وتعلم الآلة المراقب. ومع ذلك، يمكن أن تكون هذه الأساليب التقليدية غير فعالة، خاصة عند التعامل مع مجموعات بيانات كبيرة. تقدم التطورات الأخيرة في التعلم “القليل من الأمثلة” و”بدون أمثلة” من خلال نماذج اللغة الكبيرة التوليدية (LLMs) مثل GPT-4 وClaude 2 بديلاً واعدًا، مما قد يسهل التحليلات الموضوعية للمحتوى المتعلق بالصحة.

تهدف الورقة إلى تقييم فعالية نماذج LLMs في إجراء اختيار نماذج الموضوع والتحليل الموضوعي الاستقرائي للمنشورات على وسائل التواصل الاجتماعي التي تركز على بلاغة اللقاح. من خلال مقارنة مخرجات ثلاثة نماذج LLMs مختلفة مع التحليلات البشرية السابقة، تسعى الدراسة لتحديد ما إذا كانت نماذج LLMs يمكن أن تحقق نتائج قابلة للمقارنة في تحديد الموضوعات وترميز المحتوى. يفترض المؤلفون أن نماذج LLMs ستحدد موضوعات وموضوعات مشابهة لتلك التي حددها البشر، وستظهر تباينًا في الأداء، وستظهر موثوقية عبر التحليلات المتكررة. هذه الدراسة مهمة لأنها تعالج الفجوة في التحقق من نماذج LLMs لتحليل الموضوعات المتعلقة بالصحة، مما قد يسهل استراتيجيات مراقبة الصحة العامة والاستجابة بشكل أكثر كفاءة.

الطرق

في هذا القسم، يحدد المؤلفون المنهجيات المستخدمة لتحديد أفضل خمسة موضوعات ذات صلة تتعلق بخطاب مكافحة اللقاح على وسائل التواصل الاجتماعي، مقارنة طرق الاختيار اليدوي بتلك التي تستخدم نماذج اللغة الكبيرة (LLMs). في البداية، يصف المؤلفون دراسة سابقة حيث جمعوا 3,999,726 تغريدة من واجهة برمجة التطبيقات العامة لتويتر، مع تصفية الكلمات الرئيسية المتعلقة بكوفيد-19. بعد إزالة التكرارات، عزلوا رسائل مكافحة اللقاح وطبقوا نموذج موضوع Biterm (BTM) لتصنيف البيانات إلى 20 مجموعة. تم فحص أفضل التغريدات من هذه المجموعات يدويًا لتحديد خمسة موضوعات ذات صلة، مع التركيز على تأثير الشخصيات العامة على مواقف اللقاح.

في الدراسة الحالية، يكرر المؤلفون هذه العملية باستخدام نماذج LLMs، وبالتحديد نماذج GPT-4 وClaude، لترتيب مجموعات BTM بناءً على الصلة. أعدوا نفس مجموعة البيانات وطلبوا من نماذج LLMs تحديد أفضل خمسة موضوعات، ثم قارنوا النتائج بالموضوعات المختارة يدويًا من الدراسة السابقة. تم إجراء تقييمات إحصائية باستخدام التوزيع الفائق لتقييم احتمال التداخل بين اختيارات LLM والاختيارات البشرية. بالإضافة إلى ذلك، استخدم المؤلفون التحليل الموضوعي الاستقرائي لاشتقاق الموضوعات من نفس مجموعة التغريدات، متناقضين نتائج تحليل LLM مع تلك الناتجة عن التحليل البشري التقليدي وتقنيات نمذجة الموضوعات الأخرى، مثل تخصيص ديريشليت الكامن (LDA) وBERTopic. شملت المنهجيات أيضًا تقييمات للهلاوس في مخرجات LLM ووضوح الموضوعات المستمدة من كل من نماذج LLM ونماذج الموضوعات، مما يبرز الاختلافات الدقيقة في تفسير الموضوعات وإمكانية نماذج LLM لتبسيط عملية التحليل.

النتائج

يقدم قسم النتائج من الدراسة تحليلًا مقارنًا لملاءمة الموضوع والتحليل الموضوعي بين الاختيارات البشرية والمخرجات التي تولدها نماذج اللغة الكبيرة (LLMs). تم التعرف على أفضل خمسة موضوعات حددها المحللون البشر أيضًا من قبل نماذج LLMs، حيث أظهر GPT أعلى دقة، يليه Claude 1. قامت كل LLM بترتيب الموضوعات العشرين، والتصنيفات مفصلة في الجدول 1. بالإضافة إلى ذلك، أنتجت نماذج LLMs تحليلات موضوعية بنجاح، مع تنسيق المخرجات كما هو مطلوب، على الرغم من أن درجة التوافق مع الموضوعات المحددة من قبل البشر كانت متفاوتة. ومن الجدير بالذكر أن Claude 1 تطابق بشكل وثيق مع الموضوعات الأصلية، بينما حددت الجولة الثانية من اختبار GPT موضوعًا واحدًا فقط.

أظهر تقييم مخرجات نماذج LLMs أنها نادرًا ما أنتجت منشورات وهمية، حيث كان أداء GPT-4 هو الأفضل من حيث الوفاء بالمحتوى الأصلي. كانت المخرجات الموضوعية من نماذج LLMs عمومًا تعتبر معقولة، مع درجات تتراوح من 1.8 إلى 2.8، مما يشير إلى تباين في الجودة. بالمقابل، فشلت نماذج الموضوعات مثل LDA وBERTopic في مطابقة أي من الموضوعات الأصلية التي أنشأها البشر وحصلت على درجات أقل من حيث المعقولية، مما يشير إلى أن نماذج LLMs تفوقت على الأساليب التقليدية لنمذجة الموضوعات في كل من توافق الموضوع ووضوح المخرجات. بشكل عام، تسلط النتائج الضوء على إمكانيات نماذج LLMs في التحليل الموضوعي بينما تشير أيضًا إلى مجالات للتحسين في التناسق والدقة عبر جولات الاختبار المختلفة.

المناقشة

في هذه الدراسة، بحثنا في فعالية نماذج اللغة الكبيرة (LLMs) في إجراء اختيار نماذج الموضوع والتحليل الموضوعي للمنشورات المتعلقة بالصحة على وسائل التواصل الاجتماعي، لا سيما فيما يتعلق بمشاعر مكافحة اللقاح. تشير نتائجنا إلى أن نماذج LLMs يمكن أن تحدد بفعالية الموضوعات ذات الصلة، حيث كان أداء GPT-4 جيدًا بشكل خاص، حيث كان يصنف باستمرار الموضوعات الرئيسية مثل “[محترف التنس] موقف المناهض للقاح” بين أفضل خمسة. ومع ذلك، تم تجاهل بعض الموضوعات، لا سيما تلك التي تتضمن شخصيات خيالية مثل “أيمي دانكان”، غالبًا من قبل نماذج LLMs، مما يبرز قيدًا في فهمها السياقي مقارنة بالمحللين البشر.

بينما أظهرت نماذج LLMs قدرة معقولة على استنباط موضوعات مشابهة لتلك التي حددها البشر، إلا أنها لم تتطابق باستمرار مع جميع الموضوعات الأصلية، مما يشير إلى أن المبرمجين البشر قد يلتقطون تفسيرات أكثر دقة. كما كشفت الدراسة عن تباين في الأداء عبر نماذج LLMs المختلفة وجولات الاختبار، مما يشير إلى أنه بينما يمكن أن تكون نماذج LLMs أدوات فعالة للتحليل الموضوعي، إلا أنها حاليًا لا تعيد إنتاج عمق ودقة التقييمات النوعية البشرية. يجب أن تستكشف الأبحاث المستقبلية الأساليب الهجينة التي تجمع بين رؤى البشر وقدرات نماذج LLMs لتعزيز اكتشاف الموضوعات ومعالجة القيود التي لوحظت في هذه الدراسة.

Journal: JMIR Infodemiology, Volume: 4
DOI: https://doi.org/10.2196/59641
PMID: https://pubmed.ncbi.nlm.nih.gov/39207842
Publication Date: 2024-07-01
Author(s): Michael Deiner et al.
Primary Topic: Misinformation and Its Impacts

Overview

The research investigates the feasibility of using generative large language models (LLMs) for topic model selection and inductive thematic analysis of social media posts related to public health. Given the challenges of manual analysis by trained experts, the study compares the performance of three LLMs—GPT4-32K, Claudeinstant-100K, and Claude-2-100K—against human analyses from a prior study on vaccine rhetoric. The results indicate that all three LLMs effectively summarized the content and identified relevant topics, with statistical significance (P<0.001) suggesting a higher likelihood of including human-rated top content areas in their rankings. While the LLMs identified themes that were generally comparable to those recognized by human experts, variability was observed between different LLMs and across repeated analyses of the same model. Despite not consistently replicating the depth of human analysis, the themes generated by the LLMs were deemed reasonable and relevant by subject matter experts. The findings highlight the potential of LLMs for automated social media listening in public health, although further validation is necessary to enhance their accuracy and reliability in extracting nuanced themes.

Introduction

The introduction of this research paper discusses the potential of social media platforms to enhance public health understanding, particularly regarding vaccine sentiment, which is vital for disease prevention and public wellbeing. Despite the wealth of user-generated content available, analyzing this unstructured data is often labor-intensive and time-consuming, prompting researchers to employ methods like natural language processing (NLP) and supervised machine learning. However, these traditional approaches can be inefficient, especially when handling large datasets. Recent advancements in “few-shot” and “zero-shot” learning through generative large language models (LLMs) like GPT-4 and Claude 2 present a promising alternative, potentially streamlining thematic analyses of health-related content.

The paper aims to evaluate the effectiveness of LLMs in conducting topic model selection and inductive thematic analysis of social media posts focused on vaccine rhetoric. By comparing the outputs of three different LLMs against previous human analyses, the study seeks to determine whether LLMs can achieve comparable results in thematic identification and content coding. The authors hypothesize that LLMs will identify similar topics and themes as humans, exhibit variability in performance, and demonstrate reliability across repeated analyses. This research is significant as it addresses the gap in validating LLMs for health-related thematic analysis, which could facilitate more efficient public health monitoring and response strategies.

Methods

In this section, the authors outline the methodologies employed to identify the top five relevant topics related to anti-vaccination discourse on social media, comparing manual selection methods with those utilizing large language models (LLMs). Initially, the authors describe a prior study where they collected 3,999,726 tweets from Twitter’s public API, filtering for COVID-19-related keywords. After removing duplicates, they isolated anti-vaccine messages and applied the Biterm Topic Model (BTM) to categorize the data into 20 clusters. The top tweets from these clusters were manually screened to identify five relevant topics, focusing on public figures’ influence on vaccination attitudes.

In the current study, the authors replicate this process using LLMs, specifically GPT-4 and Claude models, to rank the BTM groups based on relevance. They prepared the same dataset and prompted the LLMs to identify the top five topics, subsequently comparing the results to the manually selected topics from the previous study. Statistical assessments were conducted using hypergeometric distribution to evaluate the likelihood of overlap between LLM and human selections. Additionally, the authors employed inductive thematic analysis to derive themes from the same set of tweets, contrasting the outcomes of LLM analysis with those from traditional human analysis and other topic modeling techniques, such as Latent Dirichlet Allocation (LDA) and BERTopic. The methodologies also included assessments of hallucinations in LLM outputs and the clarity of themes derived from both LLMs and topic models, highlighting the nuanced differences in thematic interpretation and the potential for LLMs to streamline the analysis process.

Results

The results section of the study presents a comparative analysis of topic relevance and thematic analysis between human selections and outputs generated by large language models (LLMs). The top five topics identified by human analysts were also recognized by the LLMs, with GPT demonstrating the highest accuracy, followed by Claude 1. Each LLM ranked the 20 topics, and the rankings are detailed in Table 1. Additionally, LLMs successfully produced thematic analyses, with outputs formatted as requested, though the degree of alignment with human-defined themes varied. Notably, Claude 1 closely matched the original themes, while GPT’s second test run identified only one theme.

The assessment of LLMs’ outputs revealed that they rarely generated phantom posts, with GPT-4 performing best in terms of fidelity to original content. The thematic outputs from LLMs were generally deemed reasonable, with scores ranging from 1.8 to 2.8, indicating variability in quality. In contrast, topic models like LDA and BERTopic failed to match any of the original human-induced themes and received lower reasonableness scores, suggesting that LLMs outperformed traditional topic modeling approaches in both thematic alignment and clarity of output. Overall, the findings highlight the potential of LLMs in thematic analysis while also indicating areas for improvement in consistency and accuracy across different test runs.

Discussion

In this study, we investigated the efficacy of large language models (LLMs) in conducting topic model selection and thematic analysis of health-related social media posts, specifically regarding anti-vaccination sentiments. Our findings indicate that LLMs can effectively identify relevant topics, with GPT-4 performing particularly well, consistently ranking key topics such as “[Tennis Pro] antivaxxer stance” among the top five. However, certain themes, particularly those involving fictional characters like “Amy Duncan,” were often disregarded by LLMs, highlighting a limitation in their contextual understanding compared to human analysts.

While LLMs demonstrated a reasonable ability to induce themes similar to those identified by humans, they did not consistently match all original themes, suggesting that human coders may capture more nuanced interpretations. The study also revealed variability in performance across different LLMs and test runs, indicating that while LLMs can serve as efficient tools for thematic analysis, they currently do not replicate the depth and precision of human qualitative assessments. Future research should explore hybrid approaches that combine human insights with LLM capabilities to enhance thematic detection and address the limitations observed in this study.

شارك: