نموذج تعلم جماعي مختلط لتوقع أسعار النفط الخام
A blending ensemble learning model for crude oil price forecasting

المجلة: Annals of Operations Research، المجلد: 353، العدد: 2
DOI: https://doi.org/10.1007/s10479-023-05810-8
تاريخ النشر: 2024-01-25
المؤلف: Mahmudul Hasan وآخرون
الموضوع الرئيسي: ديناميات السوق والتقلبات

نظرة عامة

في هذه الدراسة، يقترح المؤلفون نموذج تعلم جماعي مبتكر يمزج بين تقنيات الانحدار في تعلم الآلة المتعددة—على وجه التحديد، انحدار الجار الأقرب، أشجار الانحدار، الانحدار الخطي، الانحدار الجبهي، وانحدار الدعم المتجه—لتحسين توقعات أسعار النفط الخام. يتم التحقق من صحة النموذج باستخدام بيانات أسعار النفط الخام برنت وWTI عبر ترددات زمنية مختلفة، ويتم مقارنته بأساليب فردية وجماعية معروفة، مثل الانحدار اللغوي وغابة عشوائية. تشير النتائج إلى أن النموذج المدمج المقترح، المعروف باسم LKDSR، يتفوق على الأساليب الحالية من حيث دقة التوقعات ويظهر متانة عبر درجات مختلفة من السلاسل الزمنية، وهو أمر حاسم لقرارات الصناعة.

تشير النتائج إلى أن نموذج LKDSR لا يوفر فقط أداءً محسنًا في التوقعات، بل يقدم أيضًا رؤى قيمة للمساهمين الذين يسعون لاستراتيجيات استثمار مثلى خلال تقلبات سوق النفط. ومع ذلك، تعترف الدراسة بالقيود، بما في ذلك تركيزها فقط على سلاسل زمنية لأسعار النفط الخام أحادية المتغير دون النظر في العوامل الخارجية مثل المتغيرات البيئية أو المؤشرات الاقتصادية الكلية. تشمل اتجاهات البحث المستقبلية دمج هذه المحددات الإضافية واستكشاف قابلية تطبيق النموذج على سلع وأسواق مالية أخرى. كما يبرز المؤلفون إمكانية تعزيز قدرات النموذج في التوقعات من خلال دمج مؤشرات شعور السوق واستخدام تقنيات تعلم الآلة الأوسع في الدراسات اللاحقة.

مقدمة

تسلط مقدمة الورقة الضوء على الدور الحاسم للنفط الخام كمصدر رئيسي للطاقة، حيث يمثل نحو ثلث استهلاك الطاقة العالمي. تناقش الآثار الاقتصادية لتقلبات أسعار النفط الخام، مشيرة إلى أن ارتفاع الأسعار يمكن أن يعيق النمو الاقتصادي ويساهم في التضخم، بينما قد تؤثر الأسعار المنخفضة سلبًا على الدول المصدرة للنفط. إن تقلب أسعار النفط الخام، الذي يتأثر بعوامل متعددة مثل ديناميات العرض والطلب، والأحداث الجيوسياسية، ومشاعر السوق، يتطلب طرق توقع فعالة.

يستعرض المؤلفون الأساليب الحالية لتوقع أسعار النفط، بما في ذلك الأساليب الإحصائية والاقتصادية وتقنيات تعلم الآلة، مع التركيز بشكل خاص على أساليب التعلم الجماعي. على الرغم من التقدم، يلاحظون أنه لم تتفوق أي طريقة واحدة باستمرار على الأخرى في دقة التوقعات على المدى الطويل بسبب التقلبات الكامنة في أسعار النفط. لمعالجة هذه الفجوة، تقترح الورقة نموذج تعلم جماعي مبتكر، يُعرف باسم LKDSR، والذي يدمج خمسة متنبئين متنوعين: انحدار الجار الأقرب، الانحدار الخطي، شجرة الانحدار، انحدار الدعم المتجه، والانحدار الجبهي. يهدف هذا النموذج إلى تعزيز دقة التوقعات لكل من التوقعات قصيرة ومتوسطة المدى لأسعار النفط الخام برنت وWTI، مع الإشارة إلى أن النتائج الأولية تشير إلى تفوقه على النماذج الحالية كما تم التحقق منه من خلال اختبار ديبولد ماريانو. توضح الورقة هيكلها، مع تفاصيل الأقسام اللاحقة حول مراجعة الأدبيات، وصف البيانات، المنهجية، النتائج، المناقشة، والاستنتاجات.

طرق

في هذه الدراسة، تم استخدام نموذج LKDSR لتعلم الجماعي لتوقع أسعار النفط الخام على مقاييس يومية وأسبوعية وشهرية. تكمن قوة النموذج في كفاءته الحاسوبية وتنوع متنبئيه الأساسيين، مما يسهل نمذجة أنماط أسعار النفط المختلفة. تم جمع بيانات من النفط الخام برنت وWTI ومعالجتها مسبقًا لتحسين الأداء، مع تحويل البيانات اليومية إلى تنسيقات أسبوعية وشهرية. بعد اتباع منهجية تحقق تسلسلية، تم تقسيم مجموعة البيانات إلى مجموعات تدريب واختبار بنسبة 4:1. تم تقييم أداء النموذج باستخدام مقاييس خطأ متعددة، وتم التحقق من الأهمية الإحصائية لنموذج LKDSR من خلال اختبار ديبولد-ماريانو (DM).

استخدمت الدراسة كل من خوارزميات تعلم الآلة الفردية والجماعية لتوقع الأسعار، مقارنة نموذج LKDSR بأساليب معروفة مثل Lasso، انحدار الدعم المتجه (SVR)، AdaBoost، أشجار الانحدار، LightGBM (LGB)، CatBoost، والغابة العشوائية. تم إجراء ضبط المعلمات باستخدام نهج البحث الشبكي، الذي قيم بشكل منهجي تركيبات المعلمات لتحديد التكوين الأمثل بناءً على مقاييس الأداء. تم تنفيذ جميع التجارب على محطة عمل مخصصة ويمكن أيضًا تشغيلها على Google Colab، مما يبرز إمكانية الوصول إلى المنهجية. يُلاحظ أن طريقة البحث الشبكي، على الرغم من فعاليتها، تتطلب كثافة حاسوبية، خاصة مع مساحات المعلمات الكبيرة.

نتائج

تشير نتائج توقع أسعار النفط الخام برنت إلى أن نموذج LKDSR المدمج المقترح تفوق على طرق تعلم الآلة الأخرى عبر آفاق التوقع اليومية والأسبوعية والشهرية. على وجه الخصوص، حقق نموذج LKDSR $R^2$ ملحوظًا بنسبة 99% للتوقعات اليومية، مما يظهر دقة متفوقة مع أخطاء مطلقة ونسبية أقل مقارنة بالنماذج البديلة. على الرغم من أن النماذج المعتمدة على SVR أظهرت عيوبًا كبيرة—إما بتقدير الأسعار بشكل مبالغ فيه أو تقديرها بشكل ناقص—إلا أن نموذج LKDSR قدم باستمرار توقعات أكثر موثوقية، كما هو موضح في المنحنيات الفعلية مقابل المتوقعة. بالنسبة للتوقعات الأسبوعية، حافظ نموذج LKDSR على تفوقه، بينما بالنسبة للتوقعات الشهرية، لا يزال يقدم أفضل أداء من حيث متوسط الخطأ المطلق (MAE)، على الرغم من انخفاض الدقة.

في سياق توقع أسعار النفط الخام WTI، تفوق نموذج LKDSR أيضًا في التوقعات اليومية والأسبوعية وتفوق بشكل ملحوظ على النماذج الأخرى في التوقعات الشهرية أيضًا. كشفت التحليلات أن نموذج LKDSR المدمج أظهر أداءً قويًا عبر جميع آفاق التوقع، مع تأكيد الفروق الكبيرة من خلال اختبارات ديبولد-ماريانو (DM) الإحصائية. بالإضافة إلى ذلك، قارنت الدراسة النهج المدمج باستراتيجيات جماعية أخرى، مثل التجميع والتصويت بالأغلبية، ووجدت أن نموذج LKDSR تجاوز هذه البدائل باستمرار في الأداء. وهذا يشير إلى أنه بينما يقدم التجميع تعقيدًا ومرونة، فإن الطريقة المدمجة توفر حلاً أكثر كفاءة وفعالية لتوقع أسعار النفط الخام، خاصة عندما تكون الموارد الحاسوبية اعتبارًا.

مناقشة

تستعرض قسم المناقشة في الورقة التقدمات الأخيرة في توقع أسعار النفط الخام، مع التركيز بشكل خاص على تقنيات تعلم الآلة. تسلط الضوء على فعالية نماذج متنوعة، بما في ذلك الشبكات العصبية السطحية، وهياكل التعلم العميق مثل LSTM وBi-LSTM، والنهج الهجينة التي تجمع بين الأساليب الإحصائية التقليدية مع تعلم الآلة. من الجدير بالذكر أن النماذج الجماعية أظهرت أداءً متفوقًا في التقاط تعقيدات ديناميات أسعار النفط، حيث كانت التجمعات غير المتجانسة فعالة بشكل خاص بسبب قدرتها على دمج نماذج الانحدار المتنوعة، مما يعزز المتانة ضد تقلبات السوق.

كما يتناول القسم البيانات المستخدمة في التحليل، وتحديدًا أسعار النفط الخام برنت وWTI، التي تمتد من 1987 إلى 2022، ويناقش خطوات المعالجة المسبقة المتخذة، مثل التطبيع والتعامل مع القيم المفقودة. علاوة على ذلك، يقدم عدة خوارزميات تعلم الآلة المستخدمة في الدراسة، بما في ذلك الانحدار اللغوي، الغابة العشوائية، وآلة تعزيز التدرج الخفيف (LGBM)، كل منها له مزايا مميزة من حيث الدقة وكفاءة الحوسبة. يهدف نموذج تعلم الجماعي المدمج المقترح، الذي يدمج عدة خوارزميات، إلى تحسين دقة التوقع من خلال تقليل تباين البيانات ومعالجة الإفراط في التخصيص، مما يعزز موثوقية توقعات أسعار النفط الخام.

Journal: Annals of Operations Research, Volume: 353, Issue: 2
DOI: https://doi.org/10.1007/s10479-023-05810-8
Publication Date: 2024-01-25
Author(s): Mahmudul Hasan et al.
Primary Topic: Market Dynamics and Volatility

Overview

In this study, the authors propose a novel blending ensemble learning model that integrates multiple machine learning regression techniques—specifically, k-nearest neighbor regression, regression trees, linear regression, ridge regression, and support vector regression—to enhance the forecasting of crude oil prices. The model is validated using Brent and WTI crude oil price data across various time series frequencies and is benchmarked against established individual and ensemble methods, such as lasso regression and random forest. The results indicate that the proposed blending model, referred to as LKDSR, outperforms existing methods in terms of forecasting accuracy and demonstrates robustness across different time series granularities, which is crucial for industry decision-making.

The findings suggest that the LKDSR model not only provides improved forecasting performance but also offers valuable insights for stakeholders seeking optimal investment strategies during oil market fluctuations. However, the study acknowledges limitations, including its focus solely on univariate crude oil price time series without considering external factors such as environmental variables or macroeconomic indicators. Future research directions include the incorporation of these additional determinants and the exploration of the model’s applicability to other commodities and financial markets. The authors also highlight the potential for enhancing the model’s forecasting capabilities by integrating market sentiment indicators and utilizing broader machine learning techniques in subsequent studies.

Introduction

The introduction of the paper highlights the critical role of crude oil as a primary energy source, accounting for nearly one-third of global energy consumption. It discusses the economic implications of crude oil price fluctuations, noting that rising prices can hinder economic growth and contribute to inflation, while falling prices may adversely affect oil-exporting nations. The volatility of crude oil prices, influenced by various factors such as supply and demand dynamics, geopolitical events, and market sentiment, necessitates effective forecasting methods.

The authors review existing approaches to oil price forecasting, including statistical, econometric, and machine learning techniques, with a particular emphasis on ensemble learning methods. Despite the advancements, they note that no single method has consistently outperformed others in long-term forecasting accuracy due to the inherent volatility of oil prices. To address this gap, the paper proposes a novel blending ensemble learning model, referred to as LKDSR, which integrates five diverse predictors: k-nearest neighbor regression, linear regression, regression tree, support vector regression, and ridge regression. This model aims to enhance forecasting accuracy for both short-term and medium-term predictions of Brent and West Texas Intermediate (WTI) crude oil prices, with preliminary results indicating its superiority over existing models as validated by the Diebold Mariano test. The paper outlines its structure, detailing subsequent sections on literature review, data description, methodology, results, discussion, and conclusions.

Methods

In this study, the LKDSR blending ensemble learning model was employed to forecast crude oil prices on daily, weekly, and monthly scales. The model’s strengths lie in its computational efficiency and the diversity of its base predictors, which facilitate the modeling of various oil price patterns. Data from Brent and WTI crude oil were collected and pre-processed to optimize performance, with daily data converted to weekly and monthly formats. Following a sequential validation methodology, the dataset was split into training and testing subsets in a 4:1 ratio. Model performance was assessed using multiple error metrics, and the statistical significance of the LKDSR model was validated through the Diebold-Mariano (DM) test.

The study utilized both individual and ensemble machine learning algorithms for price forecasting, comparing the LKDSR model against established methods such as Lasso, Support Vector Regression (SVR), AdaBoost, regression trees, LightGBM (LGB), CatBoost, and random forest. Hyperparameter tuning was conducted using a grid search approach, which systematically evaluated combinations of hyperparameters to identify the optimal configuration based on performance metrics. All experiments were executed on a dedicated workstation and could also be run on Google Colab, highlighting the accessibility of the methodology. The grid search method, while effective, is noted for its computational intensity, especially with large parameter spaces.

Results

The results of the Brent crude oil price forecasting indicate that the proposed blending LKDSR model outperformed other machine learning methods across daily, weekly, and monthly forecasting horizons. Specifically, the LKDSR model achieved a remarkable $R^2$ of 99% for daily forecasts, demonstrating superior accuracy with lower absolute and relative errors compared to alternative models. Although the SVR-based models exhibited significant shortcomings—either overestimating or underestimating prices—the LKDSR model consistently provided more reliable predictions, as illustrated in the actual versus predicted curves. For weekly forecasts, the LKDSR model maintained its superiority, while for monthly predictions, it still performed best in terms of Mean Absolute Error (MAE), albeit with reduced accuracy.

In the context of WTI crude oil price forecasting, the LKDSR model similarly excelled for daily and weekly predictions and notably outperformed other models for monthly forecasts as well. The analysis revealed that the blending LKDSR model demonstrated robust performance across all forecasting horizons, with significant differences confirmed through the Diebold-Mariano (DM) statistical tests. Additionally, the study compared the blending approach to other ensemble strategies, such as stacking and majority voting, finding that the LKDSR model consistently surpassed these alternatives in performance. This suggests that while stacking offers complexity and versatility, the blending method provides a more efficient and effective solution for forecasting crude oil prices, particularly when computational resources are a consideration.

Discussion

The discussion section of the paper reviews recent advancements in crude oil price forecasting, particularly emphasizing machine learning techniques. It highlights the efficacy of various models, including shallow neural networks, deep learning architectures like LSTM and Bi-LSTM, and hybrid approaches that combine traditional statistical methods with machine learning. Notably, ensemble models have demonstrated superior performance in capturing the complexities of oil price dynamics, with heterogeneous ensembles being particularly effective due to their ability to incorporate diverse regression models, thus enhancing robustness against market volatility.

The section also details the data utilized for analysis, specifically Brent and WTI crude oil prices, spanning from 1987 to 2022, and discusses the preprocessing steps taken, such as normalization and handling missing values. Furthermore, it introduces several machine learning algorithms employed in the study, including Lasso regression, Random Forest, and Light Gradient Boosting Machine (LGBM), each with distinct advantages in terms of accuracy and computational efficiency. The proposed blending ensemble learning model, which integrates multiple algorithms, aims to optimize prediction accuracy by minimizing data variance and addressing overfitting, thereby enhancing the reliability of crude oil price forecasts.