DOI: https://doi.org/10.1007/s11063-024-11622-z
تاريخ النشر: 2024-05-06
المؤلف: Sameer Poongadan وآخرون
الموضوع الرئيسي: طرق التنبؤ بسوق الأسهم
نظرة عامة
تقدم هذه الدراسة نموذجًا جديدًا لتوقع السلاسل الزمنية، يُسمى نموذج ICEEMDAN-SVD-LSTM، والذي يدمج التحليل المتكامل المحسن لنمط الانحلال التجريبي مع الضوضاء التكيفية (ICEEMDAN)، تحليل القيم الفردية (SVD)، وشبكات الذاكرة طويلة وقصيرة المدى (LSTM). تم تصميم هذا النموذج خصيصًا لتحليل البيانات غير الخطية وغير الثابتة من خلال إطار عمل ثلاثي المستويات: مستوى ICEEMDAN يقوم بتفكيك السلسلة الزمنية إلى وظائف نمط جوهرية (IMFs) وبقايا؛ مستوى SVD يقوم بإزالة الضوضاء من كل مكون IMF والبقايا؛ ومستوى LSTM يتوقع المكونات الناتجة والبقايا. يتم تحقيق التوقع النهائي للبيانات الأصلية من خلال جمع توقعات جميع المكونات والبقايا.
يتم تقييم أداء نموذج ICEEMDAN-SVD-LSTM مقارنةً بعدة نماذج موجودة، بما في ذلك LSTM، EMD-LSTM، EEMD-LSTM، CEEMDAN-LSTM، EEMD-SVD-LSTM، ICEEMDAN-LSTM، وCEEMDAN-SVD-LSTM. تُظهر التحليلات المقارنة القدرات الفائقة للنموذج المقترح في التوقع، مما يبرز فعاليته في التعامل مع بيانات السلاسل الزمنية المعقدة مقارنةً بالأساليب التقليدية للتوقع.
مقدمة
تسلط المقدمة الضوء على أهمية توقع السلاسل الزمنية غير الخطية، مشيرةً إلى أن نماذج ARIMA التقليدية غير كافية لهذا الغرض بسبب اعتمادها على الاعتماديات الخطية. تم تطوير نماذج مختلفة لتوقع السلاسل الزمنية غير الخطية، بما في ذلك تلك التي تستخدم تقنيات الذكاء الاصطناعي مثل الخوارزميات الجينية والشبكات العصبية الاصطناعية (ANNs). تشمل المساهمات الملحوظة استخدام الاتجاه الخطي لهولت ونماذج مثل BATS وTBATS لتوقع حركة مرور الويب، بالإضافة إلى تطبيق ANFIS في توقع انبعاثات أول أكسيد الكربون ومعاملات نقل الحرارة.
تناقش هذه الفقرة أيضًا تقنيات التحليل المتقدمة مثل التحليل التجريبي (EMD) وإصداراته المحسنة، التحليل التجريبي الجماعي (EEMD) والتحليل التجريبي الكامل مع الضوضاء التكيفية (CEEMDAN)، والتي تعالج قضايا مثل خلط الأنماط. يهدف إدخال ICEEMDAN إلى التخفيف من الضوضاء المتبقية والأنماط الزائفة. تؤكد الورقة على مرونة تقنيات الذكاء الحسابي (CI)، وخاصة ANNs، في نمذجة العلاقات المعقدة وغير الخطية في بيانات السلاسل الزمنية. كما تسلط الضوء على ظهور شبكات الذاكرة طويلة وقصيرة المدى (LSTM) كأداة قوية للتعامل مع الاعتماديات طويلة الأجل في البيانات المتسلسلة، مما يظهر تطبيقاتها الناجحة عبر مجالات مختلفة. تقترح الدراسة نموذج توقع جديد يدمج ICEEMDAN، تحليل القيم الفردية (SVD)، وLSTM، باستخدام مصفوفات هانكل لتفكيك البيانات وإزالة الضوضاء قبل التوقع.
الطرق
توضح فقرة المنهجية النهج المنهجي المستخدم في البحث، موضحةً التقنيات والإجراءات المحددة المستخدمة لجمع وتحليل البيانات. تؤكد على أهمية إطار عمل منظم لضمان موثوقية وصدق النتائج. يتم توضيح تطبيق هذه المنهجية من خلال دراسات حالة مختلفة أو إعدادات تجريبية، مما يظهر كيف تم تحويل المبادئ النظرية إلى تطبيقات عملية.
تظهر النتائج الرئيسية من تطبيق المنهجية رؤى مهمة حول الأسئلة البحثية المطروحة. تؤكد النتائج على فعالية الطرق المختارة في تحقيق أهداف الدراسة، مما يوفر أساسًا قويًا للاستنتاجات المستخلصة. بشكل عام، لا تعمل المنهجية فقط كدليل لتكرار الدراسة ولكنها تساهم أيضًا في الفهم الأوسع للموضوع.
النتائج
في هذه الدراسة، يُقترح نموذج هجين جديد يجمع بين ICEEMDAN، SVD، وLSTM لتوقع السلاسل الزمنية غير الخطية، ويطبق بشكل خاص على أعداد بقع الشمس لوولف. يستخدم النموذج 500 تحقيقات في ICEEMDAN لتفكيك البيانات إلى سبع وظائف نمط جوهرية (IMFs)، مع حد أقصى يبلغ 10 تكرارات غربلة. يتم استخدام تحليل القيم الفردية (SVD) على تمثيل هانكل لكل IMF لتحديد فضاءات الضوضاء، حيث يتم التعامل مع القيم الفردية التي تقل عن عتبة محددة على أنها صفر. يتم بعد ذلك توقع المكونات التي تمت إزالة الضوضاء منها باستخدام شبكة LSTM، وتظهر التوقعات المجمعة قدرات النموذج التنبؤية.
يتم تقييم أداء نموذج ICEEMDAN-SVD-LSTM مقارنةً بسبع تقنيات توقع معروفة، بما في ذلك نماذج مختلفة تعتمد على LSTM. تشير النتائج إلى أن النموذج الهجين المقترح يتفوق باستمرار على الأساليب التقليدية، مما يلتقط ديناميكيات بيانات السلاسل الزمنية غير الخطية المعقدة بشكل فعال. تدعم التمثيلات البصرية وتحليلات الأخطاء تفوق النموذج. ومع ذلك، تعترف الدراسة بالقيود المتعلقة بالكشف عن الانقطاعات الهيكلية في بيانات السلاسل الزمنية، مما يشير إلى أن العمل المستقبلي يمكن أن يدمج طرقًا إحصائية مثل اختبار تشو واختبار CUSUM لتعزيز قابلية تطبيق النموذج في السيناريوهات التي تتضمن تغييرات مفاجئة في خصائص البيانات.
المناقشة
في هذه الفقرة، تناقش الورقة تقنيات التحليل والتوقع المختلفة للسلاسل الزمنية، مع التركيز على تحليل القيم الفردية (SVD)، التحليل التجريبي (EMD)، التحليل التجريبي الجماعي (EEMD)، التحليل التجريبي الكامل مع الضوضاء التكيفية (CEEMDAN)، وCEEMDAN المحسن (ICEEMDAN). يتم استخدام SVD لفصل الإشارات النظيفة عن الضوضاء من خلال تفكيك تمثيل مصفوفة هانكل للسلسلة الزمنية إلى مصفوفات متعامدة، مما يسمح بإزالة الضوضاء بشكل فعال من خلال تحديد القيم الفردية الهامة. يتم استخدام EMD وإصداراته (EEMD، CEEMDAN، وICEEMDAN) لتفكيك السلاسل الزمنية غير الخطية وغير الثابتة إلى وظائف نمط جوهرية (IMFs) وبقايا، مما يسهل استخراج الأنماط الأساسية.
يجمع النموذج المقترح بين ICEEMDAN للتفكيك، SVD لإزالة الضوضاء، وشبكات الذاكرة طويلة وقصيرة المدى (LSTM) للتوقع. تتضمن المنهجية ثلاثة مستويات: المستوى الأول يقوم بتفكيك السلسلة الزمنية إلى IMFs وبقايا باستخدام ICEEMDAN؛ المستوى الثاني يطبق SVD لإزالة الضوضاء من هذه المكونات؛ والمستوى الثالث يستخدم LSTM لتوقع القيم المستقبلية بناءً على السلسلة التي تمت إزالة الضوضاء منها. يتم التحقق من صحة النموذج مقابل بيانات تاريخية لبقع الشمس، مما يوضح فعاليته في التقاط الديناميات الزمنية وتحسين دقة التوقع. يتم استخدام مقاييس الأداء مثل الجذر التربيعي لمتوسط الخطأ التربيعي (RMSE) ومتوسط الخطأ المطلق (MAE) لتقييم الأداء التنبؤي للنموذج مقارنةً بأساليب التحليل والتوقع الأخرى.
DOI: https://doi.org/10.1007/s11063-024-11622-z
Publication Date: 2024-05-06
Author(s): Sameer Poongadan et al.
Primary Topic: Stock Market Forecasting Methods
Overview
This study introduces a novel time series forecasting model, termed the ICEEMDAN-SVD-LSTM model, which integrates Improved Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (ICEEMDAN), Singular Value Decomposition (SVD), and Long Short-Term Memory (LSTM) networks. This model is specifically designed to analyze non-linear and non-stationary data through a three-tiered framework: the ICEEMDAN level decomposes the time series into intrinsic mode functions (IMFs) and a residue; the SVD level performs de-noising on each IMF component and the residue; and the LSTM level forecasts the resulting components and residue. The final forecast of the original data is achieved by summing the predictions of all components and the residue.
The performance of the ICEEMDAN-SVD-LSTM model is evaluated against several existing models, including LSTM, EMD-LSTM, EEMD-LSTM, CEEMDAN-LSTM, EEMD-SVD-LSTM, ICEEMDAN-LSTM, and CEEMDAN-SVD-LSTM. The comparative analysis demonstrates the superior forecasting capabilities of the proposed model, highlighting its effectiveness in handling complex time series data compared to traditional forecasting approaches.
Introduction
The introduction highlights the significance of predicting non-linear time series, noting that traditional ARIMA models are inadequate for this purpose due to their reliance on linear dependencies. Various non-linear time series prediction models have been developed, including those utilizing artificial intelligence techniques such as Genetic Algorithms and Artificial Neural Networks (ANNs). Notable contributions include the use of Holt’s linear trend and models like BATS and TBATS for web traffic forecasting, as well as the application of ANFIS in predicting carbon monoxide emissions and heat transfer coefficients.
The section further discusses advanced decomposition techniques like Empirical Mode Decomposition (EMD) and its enhanced versions, Ensemble EMD (EEMD) and Complete Ensemble EMD with Adaptive Noise (CEEMDAN), which address issues like mode mixing. The introduction of ICEEMDAN aims to mitigate residual noise and spurious modes. The paper emphasizes the flexibility of Computational Intelligence (CI) techniques, particularly ANNs, in modeling complex, non-linear relationships in time series data. It also highlights the emergence of Long Short-Term Memory (LSTM) networks as a powerful tool for handling long-term dependencies in sequential data, showcasing their successful applications across various domains. The study proposes a novel prediction model that integrates ICEEMDAN, Singular Value Decomposition (SVD), and LSTM, utilizing Hankel matrices for data decomposition and denoising before prediction.
Methods
The methodology section outlines the systematic approach employed in the research, detailing the specific techniques and procedures utilized to gather and analyze data. It emphasizes the importance of a structured framework to ensure the reliability and validity of the findings. The application of this methodology is illustrated through various case studies or experimental setups, demonstrating how the theoretical principles were translated into practical applications.
Key findings from the application of the methodology reveal significant insights into the research questions posed. The results underscore the effectiveness of the chosen methods in addressing the objectives of the study, providing a robust foundation for the conclusions drawn. Overall, the methodology not only serves as a guide for replicating the study but also contributes to the broader understanding of the subject matter.
Results
In this study, a novel hybrid model combining ICEEMDAN, SVD, and LSTM is proposed for predicting non-linear time series, specifically applied to Wolf’s Sunspot Numbers. The model utilizes 500 realizations in ICEEMDAN to decompose the data into seven Intrinsic Mode Functions (IMFs), with a maximum of 10 sifting iterations. Singular Value Decomposition (SVD) is employed on the Hankel representation of each IMF to identify noise subspaces, where singular values below a specified threshold are treated as zero. The denoised components are then forecasted using the LSTM network, and the aggregated forecasts demonstrate the model’s predictive capabilities.
The performance of the ICEEMDAN-SVD-LSTM model is benchmarked against seven established forecasting techniques, including various LSTM-based models. Results indicate that the proposed hybrid model consistently outperforms traditional methods, effectively capturing the dynamics of complex non-linear time series data. Visual representations and error analyses further substantiate the model’s superiority. However, the study acknowledges limitations regarding the detection of structural breaks in time series data, suggesting that future work could integrate statistical methods such as the Chow Test and CUSUM Test to enhance the model’s applicability in scenarios involving abrupt changes in data characteristics.
Discussion
In this section, the paper discusses various decomposition and forecasting techniques for time series analysis, focusing on Singular Value Decomposition (SVD), Empirical Mode Decomposition (EMD), Ensemble Empirical Mode Decomposition (EEMD), Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN), and Improved CEEMDAN (ICEEMDAN). SVD is utilized for separating clean signals from noise by decomposing a Hankel matrix representation of the time series into orthogonal matrices, allowing for effective denoising through the identification of significant singular values. EMD and its extensions (EEMD, CEEMDAN, and ICEEMDAN) are employed to decompose non-linear and non-stationary time series into intrinsic mode functions (IMFs) and residues, facilitating the extraction of underlying patterns.
The proposed model integrates ICEEMDAN for decomposition, SVD for denoising, and Long Short Term Memory (LSTM) networks for forecasting. The methodology involves three levels: the first level decomposes the time series into IMFs and residues using ICEEMDAN; the second level applies SVD to denoise these components; and the third level utilizes LSTM to predict future values based on the denoised series. The model is validated against historical sunspot data, demonstrating its effectiveness in capturing temporal dynamics and improving forecasting accuracy. Performance metrics such as Root Mean Squared Error (RMSE) and Mean Absolute Error (MAE) are employed to assess the model’s predictive performance compared to other decomposition and forecasting approaches.
