عمليات اتخاذ القرار ماركوف مع معايير حساسة للمخاطر: نظرة عامة
Markov decision processes with risk-sensitive criteria: an overview

المجلة: Mathematical Methods of Operations Research، المجلد: 99، العدد: 1
DOI: https://doi.org/10.1007/s00186-024-00857-0
تاريخ النشر: 2024-04-01
المؤلف: Nicole Bäuerle وآخرون
الموضوع الرئيسي: المخاطر وتحسين المحفظة الاستثمارية

نظرة عامة

يقدم هذا القسم من الورقة نظرة عامة على عمليات اتخاذ القرار ماركوف الحساسة للمخاطر (MDPs)، مع التركيز على تطبيق المعادل المؤكد المحسن كإطار لتقييم التوقع والمخاطر. تتضمن المناقشة مقاييس المخاطر المعتمدة مثل مقياس المخاطر الانتروبي وقيمة المخاطر المشروطة. التركيز هو على المشاكل الثابتة مع أفق زمني غير محدود، موضحًا الشروط اللازمة لوجود سياسات مثلى ومخططًا لإجراءات الحل المقابلة.

تستكشف الورقة أيضًا الآثار النظرية لتطبيق المعادل المؤكد المحسن بشكل متكرر وعلى المكافأة المتراكمة. تستعرض نماذج مختلفة، بما في ذلك السيناريوهات المخفضة وغير المخفضة، مما يوفر فهمًا شاملاً للمنهجيات وتطبيقات MDPs الحساسة للمخاطر في سياقات اتخاذ القرار.

مقدمة

تناقش مقدمة الورقة تطور عمليات اتخاذ القرار ماركوف (MDPs)، التي تعالج مشاكل تحسين ديناميكية عشوائية تهدف إلى تعظيم المكافآت التراكمية المتوقعة. تم صياغتها في البداية بواسطة ريتشارد بيلمان في الخمسينيات، وقد توسعت النظرية لتشمل حساسية المخاطر، مما يعكس النتائج من علم النفس والاقتصاد وعلم الأحياء. يبرز المؤلفون أن النماذج التقليدية المعتمدة على التوقع قد تكون غير كافية، داعين إلى تضمين التباين واللحظات الأعلى في معايير اتخاذ القرار. يتم توضيح ذلك من خلال استخدام دالة أسية تأخذ في الاعتبار جميع لحظات العائد العشوائي، مما يؤدي إلى دالة هدف تتضمن حساسية المخاطر من خلال معامل $\gamma$.

تهدف الورقة إلى تقديم نظرة شاملة على المعادلات المؤكدة المحسنة، وهو مفهوم قدمه بن-تال وتيبول (2007)، والذي يعمم المعادلات المؤكدة التقليدية ويشمل حالات مهمة مثل مقاييس المخاطر الانتروبية وقيمة المخاطر المشروطة. سيركز المؤلفون على المشاكل الثابتة مع آفاق زمنية غير محدودة، مناقشين السياسات المثلى وإجراءات الحل. ستستكشف الأقسام اللاحقة التطبيق المتكرر للمعادلات المؤكدة المحسنة، واستخدامها في المكافآت التراكمية، وحالة التكلفة المتوسطة الحساسة للمخاطر، مما ينتهي بملخص للتطبيقات العملية. تؤكد الورقة على منظور رياضي مع تجنب قضايا الأفق الزمني المحدود والزمن المستمر، كما تم تناولها في الأدبيات الحديثة.

مناقشة

في هذا القسم، يناقش المؤلفون مفاهيم المعادلات المؤكدة (CE) والمعادلات المؤكدة المحسنة (OCE) في سياق اتخاذ القرار تحت المخاطر. يوضحون كيف يفضل الأفراد الذين يتجنبون المخاطر النتائج المؤكدة على غير المؤكدة، حتى عندما تكون القيم المتوقعة متطابقة، باستخدام يانصيبين كمثالين. تعمل دالة المنفعة، التي تزداد بشكل صارم ومقعرة، على قياس تجنب المخاطر من خلال معامل تجنب المخاطر المطلق لأرو-برات، $\gamma(x) = -\frac{u'(x)}{u”(x)}$. يتم تعريف CE كـ $CE(X) = u^{-1}(E[u(X)])$، مما يمثل المبلغ المضمون الذي يوفر نفس المنفعة مثل النتيجة غير المؤكدة، بينما يتم صياغة OCE كـ $S_u(X) = \sup_{\eta \in \mathbb{R}} \{\eta + E[u(X – \eta)]\}$، مما يلتقط إطارًا أكثر دقة في اتخاذ القرار يتضمن الاستهلاك وتفضيلات المخاطر.

يستكشف المؤلفون أيضًا خصائص OCE، مثل التزايد، والإضافة التحولية، وعدم المساواة لجنسن، والتي تعزز مجتمعة مفهوم تجنب المخاطر في التفضيلات. يقدمون أمثلة محددة لدوال المنفعة التي تؤدي إلى مقاييس مخاطر معروفة، بما في ذلك مقياس المخاطر الانتروبي وقيمة المخاطر المشروطة (CVaR). تختتم القسم بالتأكيد على أهمية هذه المقاييس في تحسين عمليات اتخاذ القرار، لا سيما في عمليات اتخاذ القرار ماركوف، حيث يمكن دمج تفضيلات حساسة للمخاطر بشكل فعال في أساليب البرمجة الديناميكية. تبرز المناقشة التحديات الحسابية والإمكانات لاستخدام تقنيات التعلم المعزز الحديثة لمعالجة مشاكل تحسين حساسة للمخاطر.

Journal: Mathematical Methods of Operations Research, Volume: 99, Issue: 1
DOI: https://doi.org/10.1007/s00186-024-00857-0
Publication Date: 2024-04-01
Author(s): Nicole Bäuerle et al.
Primary Topic: Risk and Portfolio Optimization

Overview

This section of the paper presents an overview of risk-sensitive Markov decision processes (MDPs), emphasizing the application of the Optimized Certainty Equivalent as a framework for evaluating expectation and risk. The discussion includes established risk measures such as the entropic risk measure and Conditional Value-at-Risk. The focus is on stationary problems with an infinite time horizon, detailing the conditions necessary for the existence of optimal policies and outlining the corresponding solution procedures.

The paper further explores the theoretical implications of applying the Optimized Certainty Equivalent both recursively and to the cumulated reward. It reviews various models, including both discounted and non-discounted scenarios, thereby providing a comprehensive understanding of the methodologies and applications of risk-sensitive MDPs in decision-making contexts.

Introduction

The introduction of the paper discusses the evolution of Markov decision processes (MDPs), which address stochastic dynamic optimization problems aimed at maximizing expected cumulative rewards. Initially formulated by Richard Bellman in the 1950s, the theory has expanded to incorporate risk sensitivity, reflecting findings from psychology, economics, and biology. The authors highlight that traditional expectation-based models may be inadequate, advocating for the inclusion of variance and higher moments in decision-making criteria. This is exemplified by the use of an exponential function to account for all moments of a random payoff, leading to a target function that incorporates risk sensitivity through a parameter $\gamma$.

The paper aims to provide a comprehensive overview of Optimized Certainty Equivalents, a concept introduced by Ben-Tal and Teboulle (2007), which generalizes traditional Certainty Equivalents and includes significant cases like entropic risk measures and Conditional Value-at-Risk. The authors will focus on stationary problems with infinite time horizons, discussing optimal policies and solution procedures. Subsequent sections will explore the recursive application of Optimized Certainty Equivalents, their use in cumulative rewards, and the risk-sensitive average cost case, culminating in a summary of practical applications. The paper emphasizes a mathematical perspective while avoiding finite time horizon and continuous time issues, as addressed in recent literature.

Discussion

In this section, the authors discuss the concepts of Certainty Equivalents (CE) and Optimized Certainty Equivalents (OCE) in the context of decision-making under risk. They illustrate how risk-averse individuals prefer certain outcomes over uncertain ones, even when the expected values are identical, using two lotteries as examples. The utility function, which is strictly increasing and concave, serves to quantify risk aversion through the Arrow-Pratt absolute risk aversion coefficient, $\gamma(x) = -\frac{u'(x)}{u”(x)}$. The CE is defined as $CE(X) = u^{-1}(E[u(X)])$, representing the guaranteed amount that provides the same utility as the uncertain outcome, while the OCE is formulated as $S_u(X) = \sup_{\eta \in \mathbb{R}} \{\eta + E[u(X – \eta)]\}$, capturing a more nuanced decision-making framework that incorporates consumption and risk preferences.

The authors further explore the properties of OCE, such as monotonicity, shift additivity, and Jensen’s inequality, which collectively reinforce the notion of risk aversion in preferences. They provide specific examples of utility functions leading to well-known risk measures, including the entropic risk measure and Conditional Value-at-Risk (CVaR). The section concludes by emphasizing the importance of these measures in optimizing decision-making processes, particularly in Markov decision processes, where recursive risk-sensitive preferences can be effectively integrated into dynamic programming approaches. The discussion highlights the computational challenges and the potential for using modern reinforcement learning techniques to address risk-sensitive optimization problems.