EventChat: تنفيذ وتقييم موجه نحو المستخدم لنظام توصية محادثة مدفوع بنموذج لغوي كبير لاستكشاف الفعاليات الترفيهية في سياق الشركات الصغيرة والمتوسطة
EventChat: Implementation and user-centric evaluation of a large language model-driven conversational recommender system for exploring leisure events in an SME context

شارك:
المجلة: ACM Transactions on Recommender Systems
DOI: https://doi.org/10.1145/3803546
تاريخ النشر: 2026-03-24
المؤلف: Hannes Kunstmann وآخرون
الموضوع الرئيسي: أنظمة وتقنيات التوصية

نظرة عامة

إن دمج نماذج اللغة الكبيرة (LLMs) في أنظمة التوصية المحادثية (CRS) يمثل تقدمًا كبيرًا في قدراتها الاستراتيجية. ومع ذلك، فقد ركزت الأبحاث الحالية بشكل كبير على الجوانب التقنية لتنفيذ CRS المدفوعة بـ LLM، متجاهلة تقييمات المستخدمين والآثار الاستراتيجية على الشركات الصغيرة والمتوسطة (SMEs)، والتي تعتبر حيوية للاقتصاد العالمي. تقدم هذه الورقة تصميم وتقييم أداء CRS المدفوعة بـ LLM في سياق SME، باستخدام كل من مقاييس النظام وتعليقات المستخدم النهائي. يتم تقديم نموذج ResQue المعدل لتقييم CRS المدفوعة بـ LLM لتسهيل إمكانية التكرار في هذا المجال سريع التطور. تشير النتائج إلى دقة توصية متصورة تبلغ 85.5%، بينما تسلط الضوء أيضًا على التحديات المتعلقة بالزمن المستغرق (5.7 ثوانٍ)، وتكاليف التشغيل (وسيط قدره 0.04 دولار لكل تفاعل)، والجودة، خاصة بسبب الاعتماد على التعلم القائم على الموجهات واستخدام ChatGPT كمدرج في تقنية الجيل المعزز بالاسترجاع (RAG).

في الختام، توضح دراسة الحالة الجدوى التقنية لنشر CRS المدفوعة بـ LLM للشركات الصغيرة والمتوسطة وتقدم نموذج تقييم مختصر للبحوث المستقبلية. على الرغم من الأداء المرضي للنظام، تحدد الدراسة الحواجز الحرجة للاستخدام على المدى الطويل، بما في ذلك الزمن المستغرق، والتكاليف، وجودة التفاعل. تنبع هذه التحديات من خيارات التصميم المحددة وتؤكد على ضرورة وجود نهج معماري مبتكر في التطبيقات المدفوعة بـ LLM. يعتمد التنفيذ الناجح لمثل هذه الأنظمة في SMEs على تحقيق التوازن بين أوقات انتظار المستخدمين، والربحية، وتوقعات الميزات الواقعية، وكلها تؤثر بشكل كبير على تعقيد النظام وتكاليفه. من خلال وضع هذه الرؤى في سياق التطبيق الواقعي لـ EventChat، تساهم هذه الأبحاث في فهم CRS المدفوعة بـ LLM في مشهد SMEs.

مقدمة

تسلط مقدمة هذه الورقة البحثية الضوء على أهمية أنظمة التوصية (RS) في مساعدة اتخاذ القرار من قبل المستخدمين في ظل overload المعلومات، مشيرة إلى تأثيرها الكبير على إيرادات الشركات. غالبًا ما تحد أنظمة RS التقليدية من سيطرة المستخدم، مما يجبر المستخدمين على اختيار ردود فعلية من توصيات محددة مسبقًا. تعالج أنظمة التوصية المحادثية (CRS) هذه القيود من خلال السماح بإدخال المستخدم وتعليقاته من خلال واجهات محادثة، معززة بتقنيات التعلم الآلي. لقد حول ظهور نماذج اللغة الكبيرة (LLMs) CRS بشكل أكبر، مما يتيح تفاعلات أكثر تعقيدًا ويقلل من الحاجة إلى موارد بيانات واسعة، مما يقدم فرصة واعدة للشركات، وخاصة الشركات الصغيرة والمتوسطة (SMEs).

على الرغم من الفوائد المحتملة، لا يزال تنفيذ CRS المدفوعة بـ LLM في SMEs مليئًا بالتحديات، بما في ذلك التعقيدات التقنية، وتكاليف التشغيل، وعدم اليقين في الأداء. تركز الأطر الحالية بشكل أساسي على السيناريوهات المثالية بدلاً من القيود العملية التي تواجهها SMEs. تهدف هذه الورقة إلى سد هذه الفجوة من خلال تطوير والتحقق من صحة CRS المدفوعة بـ ChatGPT في صناعة الترفيه، مع التركيز على خيارات التصميم الفعالة من حيث الموارد المصممة خصيصًا لـ SMEs. يقترح المؤلفون نموذج ResQue المعدل لتقييم تفاعلات المستخدم، يجمع بين المقاييس الموضوعية والذاتية لتقديم تقييم شامل لأداء النظام. في النهاية، تسعى الأبحاث إلى ديمقراطية الوصول إلى CRS المدفوعة بـ LLM لـ SMEs، مع معالجة اعتبارات إدارية حاسمة وتعزيز فهم أعمق لتطبيقها في سياقات الأعمال الواقعية.

الطرق

تحدد قسم المنهجية النهج المنهجي المستخدم في البحث للتحقيق في الفرضيات المحددة. استخدمت الدراسة مزيجًا من الطرق الكمية والنوعية، بما في ذلك التحليل الإحصائي ودراسات الحالة، لجمع بيانات شاملة. تم اختيار المشاركين من خلال أخذ عينات طبقية لضمان التمثيل عبر الفئات الديموغرافية الرئيسية، وشملت جمع البيانات استبيانات ومقابلات مصممة لالتقاط كل من الرؤى العددية والسردية.

تم إجراء التحليلات الإحصائية باستخدام أدوات البرمجيات لتقييم العلاقات بين المتغيرات، مع تحديد مستويات الدلالة عند p < 0.05. تم تحليل البيانات النوعية موضوعيًا، مما سمح بتحديد الأنماط والمواضيع المتكررة التي ظهرت من ردود المشاركين. سهل هذا النهج المختلط فحصًا قويًا لأسئلة البحث، مما أدى في النهاية إلى نتائج أكثر دقة تساهم في الجسم المعرفي الحالي في هذا المجال.

النتائج

تشير نتائج تحليل المسار إلى ملاءمة نموذج مرضية، مع مؤشرات مثل $\chi^2 = 3.942$، $p = .05$، CFI = 0.948، TLI = 0.917، RMSEA = 0.081، و SRMR = 0.087. تكشف النتائج الرئيسية أن دقة التوصية المتصورة والتناسق مرتبطان إيجابيًا بالفائدة المتصورة، بينما يرتبط أداء معالجة المدخلات بزيادة التحكم، مما يعزز بدوره الفائدة المتصورة. علاوة على ذلك، تؤثر الفائدة المتصورة بشكل كبير على ثقة المستخدم ورضاهم العام، حيث يعتبر الأخير مؤشرًا قويًا على نوايا الاستخدام المستقبلية.

بشكل عام، تدعم التحليلات تضمين البنى المتعلقة بجودة المحادثة، مثل التناسق وأداء معالجة المدخلات، في نموذج ResQue المعدل لأنظمة التوصية التعاونية المدفوعة بـ LLM (CRS). يظهر النموذج استقرارًا عبر مواصفات مختلفة، مع ملاحظات بتحسينات عندما تم استبعاد العلاقات الأقل أهمية، مما يؤكد قوته وبساطته في التقاط تقييمات أنظمة التوصية العامة (Gen-RecSys). تتوفر تفاصيل إضافية، بما في ذلك معاملات المسار وتعديلات النموذج، في المواد التكميلية.

المناقشة

تسلط قسم المناقشة من الورقة الضوء على الإمكانيات الاستراتيجية لأنظمة التوصية المحادثية المدفوعة بنموذج اللغة الكبير (LLM) للشركات الصغيرة والمتوسطة (SMEs). على الرغم من الفوائد المثبتة لأنظمة التوصية (RS) في تعزيز المبيعات وتفاعل العملاء عبر مختلف الصناعات، فإن الأبحاث التجريبية التي تفحص CRS المدفوعة بـ LLM من منظور SMEs محدودة. يُعتبر دمج تكنولوجيا LLM في CRS فرصة تحويلية لـ SMEs، مما يمكّن من تحسين التخصيص وتفاعل المستخدم، وهو أمر حاسم للحفاظ على القدرة التنافسية في سوق مدفوع بالبيانات. ومع ذلك، غالبًا ما تعيق القيود الموارد SMEs من الاستفادة الفعالة من هذه التقنيات المتقدمة، مما قد يوسع الفجوة بين الشركات الصغيرة والكبيرة.

تناقش الورقة أيضًا ثلاثة أطر لتنفيذ CRS المدفوعة بـ LLM، كل منها يعالج تحديات محددة مثل إدارة استفسارات المستخدم وضمان التواصل الفعال. تقدم هذه الأطر—RecLLM، InteRecAgent، وLLMCRS—طرقًا متنوعة لتعزيز تجربة المستخدم وكفاءة التشغيل. ومع ذلك، يجادل المؤلفون بأن الأطر الحالية لا تلبي بشكل كافٍ الاحتياجات الفريدة لـ SMEs، مما يستدعي تطوير EventChat، وهو CRS مدفوع بـ LLM مصمم لتحسين استخدام الموارد مع معالجة الأهداف المزدوجة لتعزيز تجربة المستخدم وتقليل مشكلات جودة البيانات. يركز تصميم EventChat على الجدوى بدلاً من التعقيد الخوارزمي، ويشمل ميزات تسهل اكتشاف خيارات الترفيه بكفاءة بينما تظل قابلة للتكيف مع القيود التي تواجهها SMEs.

أخيرًا، تؤكد الورقة على الحاجة إلى إطار تقييم معدل لـ CRS المدفوعة بـ LLM، يتجاوز مقاييس الأداء التقليدية ليشمل تعليقات المستخدم وبيانات التفاعل الموضوعية. يهدف هذا النهج متعدد الطرق إلى التقاط تعقيدات تجارب المستخدمين مع الأنظمة المحادثة، مما يضمن أن تتمكن SMEs من تقييم تأثير تقنيات LLM على نتائج أعمالها بشكل فعال.

القيود

تسلط الأبحاث الضوء على قيود كبيرة مرتبطة بالأساليب المعتمدة على الموجهات فقط في أنظمة التوصية السياقية (CRS)، خاصة مع زيادة التعقيد السياقي. وجدت الدراسة أن الاعتماد فقط على تقنيات التعلم القائم على الموجهات، مثل التعلم في السياق (ICL) وسلسلة التفكير القليلة (CoT)، أدى إلى مشكلات جودة بسبب سوء تفسير ChatGPT أحيانًا للمعلومات السياقية. وقد تفاقم هذا بسبب الحاجة إلى موجهات أطول لالتقاط معلومات مفصلة، والتي غالبًا ما تمت معالجتها بشكل غير كافٍ. بينما يمكن أن يعزز دمج مدير حوار موحد جودة الاستجابة، إلا أنه لن يعالج المشكلة الأساسية المتمثلة في نقص المعرفة السياقية الدقيقة. وبالتالي، تشير النتائج إلى أنه بينما قد تكون الهياكل المعتمدة على الموجهات كافية لمهام التوصية الأبسط، فإنها تتطلب تعزيزًا للسيناريوهات الأكثر تعقيدًا التي تتطلب معرفة أعمق في المجال والتخصيص.

بالإضافة إلى ذلك، تعترف الدراسة بعدة قيود، بما في ذلك تركيزها السياقي على شركة ناشئة في صناعة الترفيه، مما قد يقيد إمكانية تعميم النتائج على الشركات الأكبر. كما أن حجم العينة المتواضع ونقص بيانات خلفية المستخدمين الواسعة، بسبب تحديات الامتثال لـ GDPR، تعقد أيضًا القدرة على التحكم في انحياز الاختيار. قد يؤدي استبعاد بعض البنى، على الرغم من تبريره نظريًا، إلى تجاهل عوامل تنبؤية ذات صلة في سياقات CRS الأخرى. علاوة على ذلك، يحد غياب دراسات الإزالة من فهم متانة تصميم النظام عبر نماذج اللغة الكبيرة المختلفة (LLMs). على الرغم من هذه القيود، توفر الأبحاث رؤى قيمة من العالم الحقيقي حول القيود التشغيلية التي تواجهها SMEs في نشر CRS المدفوعة بـ LLM، مما يقترح طرقًا للبحث المستقبلي لتعزيز المتانة وقابلية النقل من خلال مقارنات خاضعة للرقابة ودراسات متعددة المواقع.

Journal: ACM Transactions on Recommender Systems
DOI: https://doi.org/10.1145/3803546
Publication Date: 2026-03-24
Author(s): Hannes Kunstmann et al.
Primary Topic: Recommender Systems and Techniques

Overview

The integration of large language models (LLMs) into conversational recommender systems (CRS) marks a significant advancement in their strategic capabilities. However, existing research has largely concentrated on the technical aspects of implementing LLM-driven CRS, neglecting user evaluations and the strategic implications for small to medium enterprises (SMEs), which are vital to the global economy. This paper presents the design and performance evaluation of an LLM-driven CRS within an SME context, utilizing both system metrics and end-user feedback. The revised ResQue model for evaluating LLM-driven CRS is introduced to facilitate replicability in this rapidly evolving field. The findings indicate a perceived recommendation accuracy of 85.5%, while also highlighting challenges related to latency (5.7 seconds), operational costs (median of $0.04 per interaction), and quality, particularly due to the reliance on prompt-based learning and the use of ChatGPT as a ranker in the retrieval-augmented generation (RAG) technique.

In conclusion, the case study illustrates the technical feasibility of deploying an LLM-driven CRS for SMEs and offers a concise evaluation model for future research. Despite satisfactory system performance, the study identifies critical barriers to long-term usage, including latency, costs, and interaction quality. These challenges stem from specific design choices and underscore the necessity for innovative architectural approaches in LLM-driven applications. The successful implementation of such systems in SMEs hinges on balancing user wait times, profitability, and realistic feature expectations, all of which significantly impact system complexity and costs. By contextualizing these insights within the real-world application of EventChat, this research contributes to the understanding of LLM-driven CRS in the SME landscape.

Introduction

The introduction of this research paper highlights the significance of recommender systems (RS) in aiding user decision-making amidst information overload, noting their substantial impact on firm revenues. Traditional RS often limit user control, compelling users to reactively select from pre-defined recommendations. Conversational recommender systems (CRS) address this limitation by allowing user input and feedback through conversational interfaces, enhanced by machine learning techniques. The advent of large language models (LLMs) has further transformed CRS, enabling more sophisticated interactions and reducing the need for extensive data resources, thus presenting a promising opportunity for firms, particularly small-to-medium enterprises (SMEs).

Despite the potential benefits, the implementation of LLM-driven CRS in SMEs remains fraught with challenges, including technical complexities, operational costs, and performance uncertainties. Existing frameworks primarily focus on idealized scenarios rather than the practical constraints faced by SMEs. This paper aims to bridge this gap by developing and validating a ChatGPT-driven CRS within the leisure industry, emphasizing resource-efficient design choices tailored for SMEs. The authors propose a revised ResQue model for evaluating user interactions, combining both objective and subjective metrics to provide a comprehensive assessment of system performance. Ultimately, the research seeks to democratize access to LLM-driven CRS for SMEs, addressing critical managerial considerations and fostering a deeper understanding of their applicability in real-world business contexts.

Methods

The methodology section outlines the systematic approach employed in the research to investigate the specified hypotheses. The study utilized a combination of quantitative and qualitative methods, including statistical analysis and case studies, to gather comprehensive data. Participants were selected through stratified sampling to ensure representation across key demographics, and data collection involved surveys and interviews designed to capture both numerical and narrative insights.

Statistical analyses were performed using software tools to evaluate the relationships between variables, with significance levels set at p < 0.05. The qualitative data were analyzed thematically, allowing for the identification of recurring patterns and themes that emerged from participant responses. This mixed-methods approach facilitated a robust examination of the research questions, ultimately leading to more nuanced findings that contribute to the existing body of knowledge in the field.

Results

The path analysis results indicate a satisfactory model fit, with indices such as $\chi^2 = 3.942$, $p = .05$, CFI = 0.948, TLI = 0.917, RMSEA = 0.081, and SRMR = 0.087. Key findings reveal that self-reported Recommendation Accuracy and Consistency are positively correlated with Perceived Usefulness, while Input Processing Performance is linked to higher Control, which subsequently enhances Perceived Usefulness. Furthermore, Perceived Usefulness significantly influences user Confidence and Overall Satisfaction, with the latter being a strong predictor of Future Use intentions.

Overall, the analysis supports the inclusion of constructs related to conversational quality, such as Consistency and Input Processing Performance, in the revised ResQue model for LLM-driven Collaborative Recommendation Systems (CRS). The model demonstrates stability across various specifications, with improvements noted as less significant relationships were excluded, affirming its robustness and parsimony in capturing evaluations of General Recommendation Systems (Gen-RecSys). Additional details, including path coefficients and model modifications, are available in the supplementary materials.

Discussion

The discussion section of the paper highlights the strategic potential of Large Language Model (LLM)-driven Conversational Recommendation Systems (CRS) for Small and Medium Enterprises (SMEs). Despite the established benefits of recommendation systems (RS) in enhancing sales and customer engagement across various industries, empirical research specifically examining LLM-driven CRS from the SME perspective is limited. The integration of LLM technology into CRS is posited as a transformative opportunity for SMEs, enabling improved personalization and user interaction, which is critical for maintaining competitiveness in a data-driven market. However, resource constraints often hinder SMEs from effectively leveraging these advanced technologies, potentially widening the gap between smaller and larger firms.

The paper further discusses three frameworks for implementing LLM-driven CRS, each addressing specific challenges such as managing user inquiries and ensuring effective communication. These frameworks—RecLLM, InteRecAgent, and LLMCRS—offer varying approaches to enhance user experience and operational efficiency. However, the authors argue that existing frameworks do not adequately cater to the unique needs of SMEs, prompting the development of EventChat, a tailored LLM-driven CRS designed to optimize resource use while addressing the dual objectives of enhancing user experience and mitigating data quality issues. EventChat’s design prioritizes feasibility over algorithmic sophistication, incorporating features that facilitate efficient leisure option discovery while remaining adaptable to the constraints faced by SMEs.

Finally, the paper emphasizes the need for a revised evaluation framework for LLM-driven CRS, moving beyond traditional performance metrics to include user feedback and objective interaction data. This multi-method approach aims to capture the complexities of user experiences with conversational systems, ensuring that SMEs can effectively assess the impact of LLM technologies on their business outcomes.

Limitations

The research highlights significant limitations associated with prompt-only approaches in contextual recommendation systems (CRS), particularly as contextual complexity increases. The study found that relying solely on prompt-based learning techniques, such as in-context learning (ICL) and few-shot chain-of-thought (CoT), led to quality issues due to ChatGPT’s occasional misinterpretation of contextual information. This was exacerbated by the need for longer prompts to capture detailed information, which were often inadequately processed. While the integration of a unified dialog manager could enhance response quality, it would not address the fundamental issue of missing nuanced contextual knowledge. Thus, the findings suggest that while prompt-only architectures may suffice for simpler recommendation tasks, they require augmentation for more complex scenarios involving deeper domain knowledge and personalization.

Additionally, the study acknowledges several limitations, including its contextual focus on a startup in the leisure industry, which may restrict generalizability to larger firms. The modest sample size and lack of extensive user background data, due to GDPR compliance challenges, further complicate the ability to control for selection bias. The exclusion of certain constructs, while theoretically justified, may overlook relevant predictive factors in other CRS contexts. Furthermore, the absence of ablation studies limits the understanding of the robustness of the system design across different large language models (LLMs). Despite these limitations, the research provides valuable real-world insights into the operational constraints faced by SMEs in deploying LLM-driven CRS, suggesting avenues for future research to enhance robustness and transferability through controlled comparisons and multi-site studies.

شارك: