تخصيص محفظة متعددة الفترات عالية الأبعاد باستخدام التعلم العميق المعزز
High-dimensional multi-period portfolio allocation using deep reinforcement learning

شارك:
المجلة: International Review of Economics & Finance، المجلد: 98
DOI: https://doi.org/10.1016/j.iref.2025.103996
تاريخ النشر: 2025-02-27
المؤلف: Yifu Jiang وآخرون
الموضوع الرئيسي: المخاطر وتحسين المحفظة الاستثمارية

نظرة عامة

تقدم هذه الورقة البحثية استراتيجية استثمار جديدة تستخدم التعلم العميق المعزز (DRL) لتخصيص المحفظة على المدى الطويل، مع معالجة تكاليف المعاملات والمخاطر. الإطار المقترح يدمج الشبكات العصبية التلافيفية (CNN) لنمذجة ديناميكيات الأسعار وWaveNet لالتقاط اعتمادات الأصول، كل ذلك ضمن هيكل معادلة بيلمان متعددة الفترات. تتيح هذه الطريقة تحسينًا ديناميكيًا عبر مجموعة متنوعة من الأصول المترابطة ذات المخاطر. تظهر الاختبارات التجريبية عبر فترات احتفاظ مختلفة، ومستويات تجنب المخاطر، وتكاليف المعاملات، ومؤشرات مالية فعالية الاستراتيجية وتفوقها على طرق التحسين التقليدية وغيرها من أساليب التعلم الآلي.

في الختام، تقدم الدراسة طريقة متقدمة لاختيار المحفظة متعددة الفترات تستفيد من DRL لاتخاذ القرار، جنبًا إلى جنب مع CNN وWaveNet لديناميكيات أسعار الأصول والاعتمادات المتبادلة. تشير النتائج إلى وجود علاقة أحادية بين الربحية المعدلة حسب المخاطر وأفق الاستثمار، خاصة تحت مستويات منخفضة من تجنب المخاطر وتكاليف المعاملات. ومع ذلك، فإن زيادة تجنب المخاطر تقلل من أداء المحافظ طويلة الأجل، بينما تؤثر تكاليف المعاملات سلبًا على العوائد السنوية الصافية، خاصة للاستثمارات طويلة الأجل. تتفوق الاستراتيجية المقترحة باستمرار على الطرق المنافسة في بناء محافظ مثالية، مما يظهر قوتها عبر عوامل مختلفة، بما في ذلك كمية الأصول، ومستويات تجنب المخاطر، وتكاليف المعاملات.

مقدمة

ت outlines مقدمة هذه الورقة البحثية تطور إدارة المحافظ، مع التأكيد على الانتقال من نماذج العائد الفردية التقليدية، مثل تحسين متوسط-تباين ماركويتز، إلى استراتيجيات متعددة الفترات أكثر ديناميكية. بينما كان إطار عمل ماركويتز أساسيًا، إلا أنه محدود في معالجة آفاق الاستثمار طويلة الأجل وتعقيدات ظروف السوق المتغيرة. عمل رائد من قبل ميرتون وسع اختيار المحفظة إلى الوقت المستمر، مما يبرز ضرورة تخصيص الأصول الديناميكي الذي يتكيف مع مرور الوقت. ومع ذلك، فإن تحدي اشتقاق حلول مغلقة لنمذجة التحسين متعددة الفترات قد أعاق الاعتماد الواسع لهذه النماذج المتقدمة.

لقد بدأت التقدمات الأخيرة في الطرق العددية والتعلم الآلي، وخاصة التعلم العميق المعزز (DRL)، في معالجة هذه القيود. تتيح أطر DRL تطوير استراتيجيات محافظ قابلة للتكيف يمكنها تحسين تخصيص الأصول في الوقت الحقيقي، مع استيعاب البيانات عالية الأبعاد وبيئات السوق المعقدة. تقترح الورقة نموذج اختيار محفظة متعددة الفترات جديد يدمج DRL مع معادلة بيلمان متعددة الفترات ضمن إطار عملية قرار ماركوف (MDP). تأخذ هذه الطريقة في الاعتبار مستويات مختلفة من تجنب المخاطر وقيود المحفظة، كما تستفيد من تقنيات التعلم الآلي، مثل الشبكات العصبية التلافيفية وWaveNet، لنمذجة ديناميكيات أسعار الأصول والاعتمادات العرضية. تظهر النتائج التجريبية فعالية هذه الاستراتيجية المتقدمة عبر آفاق استثمار مختلفة وملفات مخاطر، مما يمهد الطريق لتحليل شامل في الأقسام اللاحقة من الورقة.

طرق

في هذا القسم، ي outlines المؤلفون منهجيتهم لبناء محافظ استثمارية باستخدام مؤشرات مالية مختلفة، تحديدًا S&P100 وDJIA للولايات المتحدة، وS&P/TSX Composite Index لكندا، على مجموعة بيانات تمتد من 1 أبريل 2010 إلى 7 ديسمبر 2023. يتم تقسيم البيانات إلى مجموعة تدريب (من 1 أبريل 2010 إلى 31 ديسمبر 2018) ومجموعة اختبار (من 2 يناير 2019 إلى 7 ديسمبر 2023) لتقييم أداء استراتيجيات تخصيص المحفظة المقترحة تحت آفاق استثمار مختلفة. تهدف الدراسة إلى تقييم قوة المنهج المقترح ضد عوامل مثل تجنب المخاطر، وقيود وزن المحفظة، وتكاليف المعاملات، بينما تقارن أيضًا مع طرق معتمدة، بما في ذلك تقنيات التعلم الآلي واستراتيجية المحفظة التقليدية ذات الوزن المتساوي (EW).

يقترح المؤلفون عدة طرق لبناء المحفظة المثلى، بما في ذلك طريقتهم المتقدمة القائمة على التعلم العميق المعزز (DRL) متعددة الفترات (MP-Adv-DRL-Cor)، وطريقة اختيار المحفظة الحساسة للتكلفة متعددة الفترات (MP-CS-PPN-Cor)، وطريقة قائمة على DPG متعددة الفترات (MP-DPG). تتضمن كل طريقة تكاليف المعاملات في عملية التحسين، كما هو محدد بواسطة دالة مكافأة هدف محددة. بالإضافة إلى ذلك، يؤكد المؤلفون على أهمية اختيار معلمات فرعية مناسبة لنماذج التعلم الآلي، مثل معدلات التعلم وبنية الشبكات العصبية، لتعزيز كفاءة التعلم. يحددون قيودًا على أوزان المحفظة ومعدلات الخصم بناءً على الأدبيات المعتمدة لمنع الرفع المفرط وضمان استقرار النموذج.

مناقشة

ت outlines قسم المناقشة في الورقة إطارًا نظريًا لتحسين المحفظة متعددة الفترات، مع التأكيد على استراتيجية إعادة التوازن بدلاً من نهج الشراء والاحتفاظ. يحدد المؤلفون مشكلة تحسين ديناميكية توزع رأس المال بين $N$ من الأصول المترابطة على مدى أفق تخطيط من $h$ فترة. يتم تقييم أداء المحفظة من خلال سلسلة من الصيغ الرياضية، بما في ذلك حساب عوائد الأصول، وقيمة المحفظة، وتكاليف المعاملات. تتضمن دالة المنفعة إطار متوسط-تباين، حيث تتأثر تفضيلات المستثمر بمعامل تجنب المخاطر $\lambda$ ومعامل تفضيل الوقت $\gamma$. الهدف هو تعظيم المنفعة المتوقعة على المدى الطويل عبر أفق الاستثمار مع الالتزام بقيود مثل الميزانية، والتداول، وقيود الصندوق.

لمعالجة تعقيدات هذه المشكلة في التحسين، يقترح المؤلفون نهج التعلم العميق المعزز (DRL) الذي يدمج الشبكات العصبية التلافيفية (CNN) وWaveNet لالتقاط ديناميكيات أسعار الأصول والاعتمادات بين الأصول. يتم هيكلة إطار DRL كعملية قرار ماركوف (MDP)، حيث يتفاعل المستثمر (الوكيل) مع بيئة السوق لتحسين أوزان المحفظة مع مرور الوقت. تبرز الورقة أهمية تحقيق التوازن بين المخاطر والعوائد مع تضمين عقوبات لانتهاك القيود، بهدف تعزيز أداء المحفظة من خلال استراتيجيات التعلم التكيفية. تشير قسم التطبيق التجريبي إلى أن الطريقة المقترحة تظهر وعدًا في تحسين عوائد المحفظة، خاصة على مدى آفاق استثمار ممتدة، مع الأخذ في الاعتبار المخاطر المرتبطة.

Journal: International Review of Economics & Finance, Volume: 98
DOI: https://doi.org/10.1016/j.iref.2025.103996
Publication Date: 2025-02-27
Author(s): Yifu Jiang et al.
Primary Topic: Risk and Portfolio Optimization

Overview

This research paper introduces a novel investment strategy utilizing deep reinforcement learning (DRL) for long-term portfolio allocation, addressing transaction costs and risk aversion. The proposed framework integrates convolutional neural networks (CNN) to model price dynamics and WaveNet to capture asset dependencies, all within a multi-period Bellman equation structure. This approach allows for dynamic optimization across a diverse array of correlated risky assets. Empirical tests across various holding periods, risk aversion levels, transaction costs, and financial indices demonstrate the strategy’s effectiveness and superiority over traditional optimization methods and other machine learning approaches.

In conclusion, the study presents a sophisticated multi-period portfolio selection method leveraging DRL for decision-making, alongside CNN and WaveNet for asset price dynamics and interdependencies. The findings indicate a monotonic relationship between risk-adjusted profitability and investment horizon, particularly under low risk aversion and transaction costs. However, increased risk aversion diminishes the performance of long-term portfolios, while transaction costs negatively impact net annual returns, particularly for long-term investments. The proposed strategy consistently outperforms competing methods in constructing optimal portfolios, showcasing robustness across various factors, including asset quantity, risk aversion levels, and transaction costs.

Introduction

The introduction of this research paper outlines the evolution of portfolio management, emphasizing the transition from traditional single-period return models, such as Markowitz’s mean-variance optimization, to more dynamic multi-period strategies. While Markowitz’s framework has been foundational, it is limited in addressing long-term investment horizons and the complexities of changing market conditions. Pioneering work by Merton extended portfolio selection into continuous time, highlighting the necessity for dynamic asset allocation that adjusts over time. However, the challenge of deriving closed-form solutions for multi-period optimization has hindered the widespread adoption of these advanced models.

Recent advancements in numerical methods and machine learning, particularly deep reinforcement learning (DRL), have begun to address these limitations. DRL frameworks allow for the development of adaptive portfolio strategies that can optimize asset allocation in real-time, accommodating high-dimensional data and complex market environments. The paper proposes a novel multi-period portfolio selection model that integrates DRL with a multi-period Bellman equation within a Markov Decision Process (MDP) framework. This approach not only considers varying levels of risk aversion and portfolio constraints but also leverages machine learning techniques, such as convolutional neural networks and WaveNet, to model asset price dynamics and cross-sectional dependencies. Empirical results demonstrate the effectiveness of this advanced portfolio strategy across different investment horizons and risk profiles, setting the stage for a comprehensive analysis in subsequent sections of the paper.

Methods

In this section, the authors outline their methodology for constructing investment portfolios using various financial indices, specifically the S&P100 and DJIA for the US, and the S&P/TSX Composite Index for Canada, over a dataset spanning from April 1, 2010, to December 7, 2023. The data is divided into a training set (April 1, 2010, to December 31, 2018) and a test set (January 2, 2019, to December 7, 2023) to evaluate the performance of their proposed portfolio allocation strategies under different investment horizons. The study aims to assess the robustness of the proposed approach against factors such as risk aversion, portfolio weight constraints, and transaction costs, while also benchmarking against established methods, including machine learning techniques and a traditional equally-weighted (EW) portfolio strategy.

The authors propose several methods for optimal portfolio construction, including their advanced multiperiod deep reinforcement learning (DRL)-based method (MP-Adv-DRL-Cor), a multi-period cost-sensitive portfolio selection method (MP-CS-PPN-Cor), and a multi-period DPG-based method (MP-DPG). Each method incorporates transaction costs into the optimization process, as defined by a specific objective reward function. Additionally, the authors emphasize the importance of selecting appropriate hyperparameters for the machine learning models, such as learning rates and the architecture of neural networks, to enhance learning efficiency. They set constraints on portfolio weights and discount rates based on established literature to prevent excessive leverage and ensure model stability.

Discussion

The discussion section of the paper outlines a theoretical framework for multi-period portfolio optimization, emphasizing a rebalancing strategy over a buy-and-hold approach. The authors define a dynamic optimization problem that allocates capital among $N$ risky assets over a planning horizon of $h$ periods. The portfolio’s performance is assessed through a series of mathematical formulations, including the calculation of asset returns, portfolio value, and transaction costs. The utility function incorporates a mean-variance framework, where the investor’s preferences are influenced by a risk-aversion parameter $\lambda$ and a time preference parameter $\gamma$. The goal is to maximize the expected long-term utility across the investment horizon while adhering to constraints such as budget, turnover, and box constraints.

To address the complexities of this optimization problem, the authors propose a deep reinforcement learning (DRL) approach that integrates convolutional neural networks (CNN) and WaveNet for capturing asset price dynamics and cross-asset dependencies. The DRL framework is structured as a Markov Decision Process (MDP), where the investor (agent) interacts with the market environment to optimize portfolio weights over time. The paper highlights the importance of balancing risk and return while incorporating penalties for violating constraints, ultimately aiming to enhance portfolio performance through adaptive learning strategies. The empirical application section indicates that the proposed method shows promise in improving portfolio returns, particularly over extended investment horizons, while also considering the associated risks.

شارك: