DOI: https://doi.org/10.1177/00131644241302721
PMID: https://pubmed.ncbi.nlm.nih.gov/39759537
تاريخ النشر: 2025-01-03
المؤلف: Duygu Koçak
الموضوع الرئيسي: المنهجيات النفسية والاختبار
نظرة عامة
تدرس هذه الدراسة فعالية ChatGPT، أداة المحادثة الذكية التي طورتها OpenAI، في إجراء تحليل البيانات من خلال تحليل العوامل الاستكشافية (EFA). قام الباحثون بإنشاء مجموعات بيانات محاكاة تحت ظروف مختلفة، بما في ذلك التوزيع الطبيعي، فئات الاستجابة، أحجام العينات، أطوال الاختبارات، تحميلات العوامل، ونماذج القياس. تم استخدام ChatGPT-4o لتحليل البيانات مرتين، بفارق أسبوع، مع مقارنة النتائج بتلك المستمدة من كود R. تم حساب مقاييس رئيسية مثل قيمة Kaiser-Meyer-Olkin (KMO)، التباين الكلي المفسر، وعدد العوامل المقدرة باستخدام معيار كايزر التجريبي، طريقة هول، ومعيار كايزر-غوتمان.
أشارت النتائج إلى أن ChatGPT أنتج نتائج متسقة مع R عبر جلستي التحليل، خاصةً للمهام الحسابية التي لا تتطلب حكمًا ذاتيًا، مثل KMO، التباين الكلي المفسر، وتحميلات العوامل. ومع ذلك، بينما كان عدد العوامل المقدرة متسقًا، لوحظت انحيازات في تحليل الهياكل متعددة الأبعاد، مما يبرز الحاجة إلى توخي الحذر من قبل الباحثين عند تفسير مثل هذه النتائج. بشكل عام، تقترح الدراسة أن ChatGPT يمكن أن يكون أداة موثوقة لبعض الجوانب الحسابية لـ EFA، ولكن قيوده في السياقات التحليلية المعقدة تتطلب اعتبارًا دقيقًا.
مقدمة
تناقش مقدمة ورقة البحث التقدم السريع والتطبيق الواسع للذكاء الاصطناعي التوليدي، مع التركيز بشكل خاص على ChatGPT، أداة معالجة اللغة الطبيعية التي طورتها OpenAI. تم إطلاق ChatGPT في 30 نوفمبر 2022، وقد شهد زيادة دراماتيكية في تفاعل المستخدمين، مع ارتفاع بنسبة 3,500% في معدلات النقر. يعتمد ChatGPT على بنية المحول المدرب مسبقًا (GPT)، وتحديداً الإصدارات GPT-3، GPT-3.5، وGPT-4، ويستخدم نهج التعلم المراقب والتعلم المعزز المعروف باسم “التعلم المعزز مع تغذية راجعة بشرية” (RLHF)، مما يعزز قدرته على إنتاج نصوص تشبه النصوص البشرية. تسلط المقدمة الضوء على قدرات هذه النماذج في مجالات متنوعة، بما في ذلك الرعاية الصحية، والمالية، والتعليم، فضلاً عن إمكانياتها في تحويل أنظمة العمل وأداء المهام المعقدة.
تناقش الورقة أيضًا القدرات التحليلية لـ ChatGPT، خاصةً مع إصدار GPT-4، الذي يمكنه إجراء تحليلات إحصائية من خلال دمج أدوات خارجية مثل Python. يشمل ذلك أداء مهام مثل اختبارات t وتحليلات العوامل، مما يؤدي إلى أتمتة العمليات التي كانت تقليديًا تتعامل معها علماء البيانات. ومع ذلك، لا تزال التحديات قائمة، خاصة فيما يتعلق بدقة وموثوقية النتائج عند التعامل مع مجموعات بيانات كبيرة أو مطالبات معقدة. تهدف الدراسة إلى تقييم أداء ChatGPT كأداة لتحليل البيانات، تحديدًا من خلال تحليل العوامل الاستكشافية (EFA)، لتقييم فعاليته ودقته في تنفيذ التحليلات الإحصائية، وهو أمر حاسم للباحثين عبر مختلف التخصصات.
النتائج
تظهر نتائج الدراسة مستوى عالٍ من التناسق بين مخرجات تحليل العوامل الاستكشافية (EFA) التي تم إنشاؤها بواسطة ChatGPT وتلك التي أنتجها R. كانت قيم KMO والتباين الكلي المفسر لشرط الـ 10 عناصر متطابقة عبر كلا النظامين، مما يشير إلى أن معلمات EFA الخاصة بـ ChatGPT تتماشى عن كثب مع تلك المستمدة من R. علاوة على ذلك، كانت التحليلات التي أجراها ChatGPT بفارق أسبوع قد أسفرت عن نتائج متسقة، مما يعزز موثوقية مخرجاته. تم حساب متوسط قوة التقدير (AEP) ونسبة الانحياز النسبي (RBR) لهياكل العوامل المختلفة، مما يكشف أن كلا الطريقتين حققت قيم AEP تتجاوز 90% وRBR أقل من 0.10، مما يشير إلى تقديرات دقيقة وغير متحيزة لعدد العوامل.
فيما يتعلق بتحميلات العوامل، وُجد أن التقديرات المقدمة من ChatGPT وR كانت متطابقة عبر جميع النماذج، بما في ذلك الهياكل ذات العامل الواحد والهياكل ثنائية الأبعاد، سواء قبل أو بعد التدوير. كما أظهرت القيم المشتركة المستمدة من التحليلات توافقًا تامًا بين الطريقتين. تشير النتائج إلى أن ChatGPT لا يؤدي فقط بشكل مشابه لـ R في EFA ولكنه أيضًا يحافظ على درجة عالية من الموثوقية والدقة في تقدير تحميلات العوامل والقيم المشتركة، مما يؤكد إمكانيته كأداة قوية للتحليل الإحصائي في بيئات البحث.
المناقشة
في هذه الدراسة، تم تقييم أداء ChatGPT في إجراء تحليل العوامل الاستكشافية (EFA) باستخدام مجموعات بيانات محاكاة. تضمنت البحث التلاعب بالمتغيرات مثل حجم العينة، تحميل العوامل، ونماذج القياس لتقييم دقة النتائج التي أنتجها ChatGPT مقارنةً بالبرامج الإحصائية التقليدية (R). أنشأت الدراسة مجموعات بيانات بحجم عينة 225، تحتوي على 10 عناصر مع تحميلات عوامل تتراوح من 0.50 إلى 0.85، وتم تحليلها تحت ظروف مختلفة. أشارت النتائج إلى أن ChatGPT أنتج باستمرار حسابات دقيقة لقيم KMO ونسب التباين المفسر، متماشية عن كثب مع تلك المستمدة من R. وهذا يشير إلى أن ChatGPT يمكن استخدامه بشكل موثوق في المهام الحسابية في EFA، شريطة أن يتم تقديم مطالبات مناسبة.
كما درست الدراسة تحديد عدد العوامل، حيث وُجد أن تقديرات ChatGPT كانت متسقة عبر التحليلات التي أجريت بفارق أسبوع. تم تقييم دقة عدد العوامل وتحميلاتها باستخدام نسبة الانحياز النسبي (RBR) ونسبة تقدير الدقة (AEP)، مما يكشف أن ChatGPT أدى بشكل مشابه لـ R، خاصة في نماذج العامل الواحد. ومع ذلك، لوحظت اختلافات طفيفة، تُعزى إلى تفسيرات محتملة خاطئة للمطالبات. بينما أظهر ChatGPT أداءً قويًا في التحليلات الإحصائية البحتة، حذرت الدراسة من أن النماذج الأكثر تعقيدًا التي تتطلب رؤى نظرية قد تستلزم حكمًا بشريًا. بشكل عام، تدعم النتائج استخدام ChatGPT كأداة قيمة للباحثين في تحليل البيانات، مع التأكيد على أهمية تقديم مطالبات واضحة والوعي بحدودها. يُشجع على إجراء أبحاث مستقبلية لاستكشاف أداء ChatGPT في تحليلات أكثر تعقيدًا وتطبيقات الذكاء الاصطناعي الأخرى في السياقات الإحصائية.
DOI: https://doi.org/10.1177/00131644241302721
PMID: https://pubmed.ncbi.nlm.nih.gov/39759537
Publication Date: 2025-01-03
Author(s): Duygu Koçak
Primary Topic: Psychometric Methodologies and Testing
Overview
This study investigates the efficacy of ChatGPT, an AI conversational tool developed by OpenAI, in performing data analysis through exploratory factor analysis (EFA). The researchers generated simulated datasets under varying conditions, including normal distribution, response categories, sample sizes, test lengths, factor loadings, and measurement models. ChatGPT-4o was employed to analyze the data twice, one week apart, with results compared to those derived from R code. Key metrics such as the Kaiser-Meyer-Olkin (KMO) value, total variance explained, and the number of factors estimated using the empirical Kaiser criterion, Hull method, and Kaiser-Guttman criterion were calculated.
The results indicated that ChatGPT produced consistent findings with R across both analysis sessions, particularly for computational tasks that do not require subjective judgment, such as KMO, total variance explained, and factor loadings. However, while the estimated number of factors was consistent, biases were noted in the analysis of multidimensional structures, highlighting the need for researchers to exercise caution when interpreting such results. Overall, the study suggests that ChatGPT can be a reliable tool for specific computational aspects of EFA, but its limitations in complex analytical contexts warrant careful consideration.
Introduction
The introduction of the research paper discusses the rapid advancement and widespread application of generative artificial intelligence (AI), particularly focusing on ChatGPT, a natural language processing tool developed by OpenAI. Launched on November 30, 2022, ChatGPT has seen a dramatic increase in user engagement, with a 3,500% rise in click-through rates. Built on the Generative Pre-trained Transformer (GPT) architecture, specifically versions GPT-3, GPT-3.5, and GPT-4, ChatGPT utilizes a supervised and reinforcement learning approach known as “Reinforcement Learning with Human Feedback” (RLHF), which enhances its ability to generate human-like text. The introduction highlights the capabilities of these models in various fields, including healthcare, finance, and education, as well as their potential for transforming work systems and performing complex tasks.
The paper also addresses the analytical capabilities of ChatGPT, particularly with the release of GPT-4, which can conduct statistical analyses by integrating external tools like Python. This includes performing tasks such as t-tests and factor analyses, thereby automating processes traditionally handled by data scientists. However, challenges remain, particularly concerning the accuracy and reliability of results when dealing with large datasets or complex prompts. The study aims to evaluate ChatGPT’s performance as a data analysis tool, specifically through exploratory factor analysis (EFA), to assess its effectiveness and accuracy in executing statistical analyses, which is crucial for researchers across various disciplines.
Results
The results of the study demonstrate a high level of consistency between the exploratory factor analysis (EFA) outputs generated by ChatGPT and those produced by R. The KMO values and total variance explained for the 10-item condition were identical across both platforms, indicating that ChatGPT’s EFA parameters align closely with those obtained from R. Furthermore, analyses conducted by ChatGPT one week apart yielded consistent results, reinforcing the reliability of its outputs. The Average Estimation Power (AEP) and Relative Bias Ratio (RBR) metrics were calculated for various factor structures, revealing that both methods achieved AEP values exceeding 90% and RBRs below 0.10, suggesting accurate and unbiased factor count estimations.
In terms of factor loadings, the estimates provided by ChatGPT and R were found to be identical across all models, including single-factor and two-dimensional structures, both before and after rotation. The communalities derived from the analyses also showed complete agreement between the two methods. The findings indicate that ChatGPT not only performs comparably to R in EFA but also maintains a high degree of reliability and accuracy in estimating factor loadings and communalities, confirming its potential as a robust tool for statistical analysis in research settings.
Discussion
In this study, the performance of ChatGPT in conducting exploratory factor analysis (EFA) was evaluated using simulated datasets. The research involved manipulating variables such as sample size, factor loading, and measurement models to assess the accuracy of findings produced by ChatGPT compared to traditional statistical software (R). The study generated datasets with a sample size of 225, containing 10 items with factor loadings ranging from 0.50 to 0.85, and analyzed them under various conditions. The results indicated that ChatGPT consistently produced accurate calculations for KMO values and explained variance ratios, aligning closely with those obtained from R. This suggests that ChatGPT can be reliably used for computational tasks in EFA, provided that appropriate prompts are given.
The study also examined the determination of the number of factors, where ChatGPT’s estimates were found to be consistent across analyses conducted one week apart. The accuracy of factor counts and loadings was evaluated using relative bias ratio (RBR) and accuracy estimation percentage (AEP), revealing that ChatGPT performed comparably to R, particularly in single-factor models. However, minor discrepancies were noted, attributed to potential misinterpretations of prompts. While ChatGPT demonstrated strong performance in purely statistical analyses, the study cautioned that more complex models requiring theoretical insight may necessitate human judgment. Overall, the findings support the use of ChatGPT as a valuable tool for researchers in data analysis, while emphasizing the importance of clear prompting and awareness of its limitations. Future research is encouraged to explore ChatGPT’s performance in more complex analyses and other AI applications in statistical contexts.
