نماذج التعلم العميق المعتمدة على اللوبي لتوقع اتجاه أسعار الأسهم: دراسة مرجعية
Lob-based deep learning models for stock price trend prediction: a benchmark study

المجلة: Artificial Intelligence Review، المجلد: 57، العدد: 5
DOI: https://doi.org/10.1007/s10462-024-10715-4
تاريخ النشر: 2024-04-13
المؤلف: Matteo Prata وآخرون
الموضوع الرئيسي: طرق التنبؤ بسوق الأسهم

نظرة عامة

تناقش هذه القسم تأثير التقدمات الأخيرة في التعلم العميق (DL) على قطاع المالية، وبشكل خاص في توقع اتجاه أسعار الأسهم (SPTP) باستخدام بيانات دفتر الطلبات المحدودة (LOB). يقوم المؤلفون بالتحقيق في قوة وعمومية خمسة عشر نموذجًا متقدمًا في التعلم العميق من خلال تطوير LOBCAST، وهو إطار عمل مفتوح المصدر مصمم لمعالجة البيانات، وتدريب النماذج، والتقييم، وتحليل الأرباح.

تشير النتائج من التجارب الواسعة إلى أن جميع النماذج التي تم فحصها تعاني من انخفاض كبير في الأداء عند مواجهة بيانات جديدة، مما يثير القلق بشأن قابليتها للتطبيق العملي في سيناريوهات السوق الحقيقية. تعتبر هذه البحث معيارًا، يسلط الضوء على كل من الإمكانيات والقيود التي تواجهها الأساليب الحالية في التعلم العميق في المالية، ويهدف إلى إبلاغ الابتكارات المستقبلية في هذا المجال.

مقدمة

تتناول مقدمة هذه الورقة البحثية التعقيدات المرتبطة بتوقع أسعار سوق الأسهم، والتي تتأثر بعوامل متنوعة مثل الظروف الاقتصادية الكلية، والأحداث الطبيعية، ومشاعر المستثمرين. تبرز أهمية النماذج الرياضية الفعالة والتقدمات الأخيرة في الذكاء الاصطناعي، وخاصة خوارزميات التعلم الآلي (ML) والتعلم العميق (DL)، التي أظهرت نتائج واعدة في البيئات المحاكاة. ومع ذلك، تشير الورقة إلى تحدٍ كبير في ترجمة هذه النجاحات إلى ظروف السوق الحقيقية، مما يشير إلى فجوة بين المحاكاة والواقع.

يقترح المؤلفون نهجًا جديدًا من خلال تقييم تقنيات التعلم العميق الحديثة لتوقع اتجاه أسعار الأسهم (SPTP) باستخدام بيانات دفتر الطلبات المحدودة (LOB)، وهو مورد حاسم للمتداولين. يقومون بتقييم قوة وعمومية نماذج التعلم الآلي والتعلم العميق المختلفة، كاشفين أنه بينما تؤدي بعض النماذج بشكل جيد على مجموعة بيانات FI-2010، فإنها تواجه صعوبة مع البيانات غير المرئية. تقدم الدراسة LOBCAST، وهو إطار عمل مفتوح المصدر لمعالجة البيانات وتقييم النماذج، وتناقش آثار نتائجهم على تحسين عمومية النماذج وتمثيل البيانات. توضح الورقة مساهماتها، بما في ذلك تقييم شامل للنماذج الحالية، وتحليل حساسية تصنيف البيانات، ومناقشة اتجاهات البحث المستقبلية.

طرق

في هذا القسم، يصف المؤلفون تنفيذهم لطريقتين تجميعتين—MAJORITY وMETALOB—لتحسين الأداء التنبؤي لـ 15 نموذجًا من نماذج التعلم العميق. تختار طريقة MAJORITY أكثر تصنيف متوقع بشكل متكرر بين المصنفات، بينما تدمج نظام وزن يعتمد على درجات F1 للمصنفات الفردية لضمان أن النماذج ذات الأداء العالي لها تأثير أكبر على القرار النهائي. بالمقابل، تستخدم طريقة METALOB شبكة عصبية متعددة الطبقات (MLP) مع طبقتين متصلتين بالكامل لتعلم كيفية دمج مخرجات المصنفات الأساسية بشكل مثالي. يتكون الإدخال إلى هذا المصنف الميتا من موتر 1D يمثل توزيع الاحتمالات على ثلاثة اتجاهات (صعود، ثابت، هبوط) لكل نموذج، مما ينتج عنه موتر يحتوي على 45 عنصرًا.

يتناول المؤلفون أيضًا التحديات المرتبطة بتوافر بيانات دفتر الطلبات المحدودة (LOB)، والتي غالبًا ما تكون مكلفة ومحدودة الوصول، خاصة من البورصات مثل NASDAQ. تعيق هذه الندرة تطوير وتطبيق خوارزميات التعلم العميق داخل مجتمع البحث. لتقييم قوة وعمومية نماذجهم، سيستخدم المؤلفون مجموعتي بيانات: FI-2010 وLOB-2021/2022، اللتين يقدمونهما في الأقسام التالية.

نقاش

في قسم النقاش، تسلط الورقة الضوء على الاهتمام المتزايد بالتعلم العميق (DL) لتوقع اتجاه أسعار الأسهم (SPTP)، مشيرة إلى دراسات متنوعة قامت بمسح النماذج والأساليب المتقدمة (SOTA). بشكل ملحوظ، تقدم Jiang (2021) وKumbure وآخرون (2022) رؤى حول هياكل الشبكات العصبية وتحليل البيانات، بينما تلخص Ozbayoglu وآخرون (2020) وSezer وآخرون (2020) خوارزميات التعلم العميق والتعلم الآلي (ML) السائدة في المالية. يؤكد المؤلفون تركيزهم الفريد على النماذج التي تستخدم بيانات دفتر الطلبات المحدودة (LOB)، مما يختلف عن الأعمال السابقة التي استعرضت بشكل أساسي الأدبيات الأوسع لتوقع أسعار الأسهم. كما يناقشون دمج تحليل المشاعر من خلال معالجة اللغة الطبيعية (NLP) لتعزيز دقة التنبؤ، كما يتضح من دراسات Shah وآخرون (2022) وAl-Alawi وAlaali (2023).

تحدد الورقة أيضًا مشكلة SPTP، مؤكدة أن بيانات LOB توفر رؤى دقيقة حول ديناميات السوق. تعرف الاتجاهات باستخدام نظام تصنيف ثلاثي (صعودي، هبوطي، مستقر) استنادًا إلى تحركات الأسعار المتوسطة، مع استخدام استراتيجية تصنيف تقلل من الضوضاء الناتجة عن تقلبات السوق. يقدم المؤلفون معيارًا شاملاً لـ 13 نموذجًا متقدمًا في التعلم العميق تم تدريبها على بيانات LOB، بما في ذلك الشبكة العصبية متعددة الطبقات (MLP) وذاكرة طويلة وقصيرة الأجل (LSTM) كخطوط أساسية. يبرزون أهمية مجموعة بيانات FI-2010 لاختبار القوة ويقدمون مجموعات بيانات LOB-2021/2022 لتقييم عمومية النماذج عبر ظروف السوق المتغيرة. تختتم الورقة بالتطرق إلى التحديات التي تطرحها تحولات توزيع البيانات في أسواق الأسهم، مما يبرز ضرورة إعادة تدريب النماذج ومراقبتها باستمرار للحفاظ على دقة التنبؤ في بيئات الإنتاج.

القيود

تسلط قيود هذه الدراسة الضوء على عدة عوامل حاسمة قد تؤثر على تفسير النتائج. أولاً، لم تكن عملية البحث عن معلمات الشبكة شاملة، مما يثير القلق بشأن احتمال عدم كفاية المعلمات المختارة في تحسين أداء النموذج. بالإضافة إلى ذلك، قيد اعتماد الدراسة على موارد حسابية محدودة مدة تدريب النماذج المعيارية إلى أسابيع بدلاً من سنوات، مما يشير إلى أن فترات التدريب الأطول قد تؤدي إلى نتائج مختلفة.

علاوة على ذلك، فإن تطبيق نماذج التعلم العميق (DL) والذكاء الاصطناعي (AI) في حل مشاكل البرمجة العشوائية والتداول (SPTP) يقدم مخاطر متأصلة. تشمل هذه المخاطر تحديات تقنية مثل تحيز البيانات، الذي يمكن أن يميل الخوارزميات التنبؤية نحو المجموعات الممثلة بشكل زائد، ونقص في قابلية تفسير النموذج مما قد يؤدي إلى اتخاذ قرارات سيئة وانخفاض الثقة في الأنظمة. كما تثار مخاوف أخلاقية، خاصة فيما يتعلق بإمكانية زيادة عدم المساواة بين المستثمرين بسبب التكاليف العالية المرتبطة بتقنيات التداول بالذكاء الاصطناعي، مما قد يعزز من عدم التوازن في السوق ويهدد الاستقرار المالي. تدعو الدراسة إلى إطار تنظيمي قوي لمعالجة هذه المخاطر مع الاستفادة من فوائد الذكاء الاصطناعي في المالية.

Journal: Artificial Intelligence Review, Volume: 57, Issue: 5
DOI: https://doi.org/10.1007/s10462-024-10715-4
Publication Date: 2024-04-13
Author(s): Matteo Prata et al.
Primary Topic: Stock Market Forecasting Methods

Overview

The section discusses the impact of recent advancements in Deep Learning (DL) on the finance sector, specifically in Stock Price Trend Prediction (SPTP) using Limit Order Book (LOB) data. The authors investigate the robustness and generalizability of fifteen state-of-the-art DL models through the development of LOBCAST, an open-source framework designed for data preprocessing, model training, evaluation, and profit analysis.

The findings from extensive experiments indicate that all examined models experience a significant decline in performance when faced with new data, raising concerns about their practical applicability in real-world market scenarios. This research serves as a benchmark, highlighting both the potential and limitations of current DL approaches in finance, and aims to inform future innovations in the field.

Introduction

The introduction of this research paper addresses the complexities involved in predicting stock market prices, which are influenced by various factors such as macroeconomic conditions, natural events, and investor sentiment. It highlights the importance of effective mathematical models and the recent advancements in artificial intelligence, particularly machine learning (ML) and deep learning (DL) algorithms, which have shown promising results in simulated environments. However, the paper notes a significant challenge in translating these successes to real market conditions, indicating a simulation-to-reality gap.

The authors propose a novel approach by benchmarking recent DL techniques for Stock Price Trend Prediction (SPTP) using Limit Order Book (LOB) data, a critical resource for traders. They evaluate the robustness and generalizability of various ML and DL models, revealing that while some models perform well on the FI-2010 dataset, they struggle with unseen data. The study introduces LOBCAST, an open-source framework for data processing and model evaluation, and discusses the implications of their findings for improving model generalizability and data representation. The paper outlines its contributions, including a comprehensive evaluation of existing models, sensitivity analysis of data labeling, and a discussion on future research directions.

Methods

In this section, the authors describe their implementation of two ensemble methods—MAJORITY and METALOB—to enhance the predictive performance of 15 deep learning models. The MAJORITY method selects the most frequently predicted class label among the classifiers, while incorporating a weighting scheme based on the F1-Scores of individual classifiers to ensure that higher-performing models have a greater influence on the final decision. Conversely, the METALOB method utilizes a multilayer perceptron (MLP) with two fully connected layers to learn how to optimally combine the outputs of the base classifiers. The input to this meta-classifier consists of a 1D tensor representing the probability distribution over three trends (up, stationary, down) for each model, resulting in a tensor with 45 elements.

The authors also address the challenges associated with the availability of limit order book (LOB) data, which is often costly and limited in access, particularly from stock exchanges like NASDAQ. This scarcity hinders the development and application of deep learning algorithms within the research community. To evaluate the robustness and generalizability of their models, the authors will utilize two datasets: FI-2010 and LOB-2021/2022, which they introduce in the subsequent sections.

Discussion

In the discussion section, the paper highlights the growing interest in deep learning (DL) for stock price trend prediction (SPTP), referencing various studies that have surveyed state-of-the-art (SOTA) models and methodologies. Notably, Jiang (2021) and Kumbure et al. (2022) provide insights into neural network structures and data analysis, while Ozbayoglu et al. (2020) and Sezer et al. (2020) summarize prevalent DL and machine learning (ML) algorithms in finance. The authors emphasize their unique focus on models utilizing limit order book (LOB) data, contrasting with prior works that primarily reviewed broader SPTP literature. They also discuss the integration of sentiment analysis through natural language processing (NLP) to enhance prediction accuracy, as evidenced by studies from Shah et al. (2022) and Al-Alawi and Alaali (2023).

The paper further delineates the SPTP problem, asserting that LOB data provides granular insights into market dynamics. It defines trends using a ternary classification system (upward, downward, stable) based on mid-price movements, employing a labelling strategy that mitigates noise from market fluctuations. The authors present a comprehensive benchmark of 13 SOTA DL models trained on LOB data, including Multilayer Perceptron (MLP) and Long-Short Term Memory (LSTM) as baselines. They highlight the significance of the FI-2010 dataset for testing robustness and introduce LOB-2021/2022 datasets to assess model generalizability across varying market conditions. The paper concludes by addressing the challenges posed by data distribution shifts in stock markets, underscoring the necessity for continual model retraining and monitoring to maintain prediction accuracy in production environments.

Limitations

The limitations of this study highlight several critical factors that may affect the interpretation of the findings. Firstly, the grid hyperparameter search employed was not exhaustive, which raises concerns about the potential inadequacy of the selected hyperparameters in optimizing model performance. Additionally, the study’s reliance on limited computational resources restricted the training duration of benchmarked models to weeks rather than years, suggesting that longer training periods could yield different outcomes.

Moreover, the application of deep learning (DL) and artificial intelligence (AI) models in solving Stochastic Programming and Trading Problems (SPTP) introduces inherent risks. These include technical challenges such as data biases, which can skew predictive algorithms towards overrepresented groups, and a lack of model explainability that may lead to poor decision-making and diminished trust in the systems. Ethical concerns also arise, particularly regarding the potential for increased inequality among investors due to the high costs associated with AI trading technologies, which could exacerbate market asymmetries and threaten financial stability. The study calls for a robust regulatory framework to address these risks while harnessing the benefits of AI in finance.