DOI: https://doi.org/10.1504/ijspm.2026.152090
تاريخ النشر: 2026-01-01
المؤلف: Zhifeng Qu
الموضوع الرئيسي: إدارة سلسلة التوريد والمخزون
نظرة عامة
تتناول هذه الدراسة الفجوة بين الأهداف التشغيلية والمالية في إدارة مخزون سلسلة التوريد من خلال تقديم إطار عمل جديد يجمع بين محاكاة الأحداث المنفصلة والتعلم العميق المعزز (DRL). النموذج المقترح، الذي يُطلق عليه نموذج تقييم الحمل المعرفي – تحسين السياسة القريبة (CLDAM-PPO)، يستخدم دالة مكافأة ذات هدفين تتضمن تكاليف المخزون التقليدية (التخزين، النقص، الطلب) جنبًا إلى جنب مع المقياس المالي لدورة تحويل النقد (CCC). تم تقييم النموذج باستخدام مجموعة بيانات دقة التنبؤ M5، حيث حقق تكلفة إجمالية قدرها 285.4 \pm 8.7$ (بالآلاف)، متفوقًا بشكل كبير على المعايير الحديثة (p < 0.01) مع الحفاظ على مستوى خدمة بنسبة 98.2% وتقليل CCC إلى 35.2 يومًا. تؤكد استنتاجات الدراسة على قدرة الإطار على سد الفجوة في اتخاذ القرارات التشغيلية والمالية في إدارة المخزون. من خلال دمج اعتبارات تكلفة التشغيل مع السيولة المالية، لا يعزز إطار CLDAM-PPO المقاييس التشغيلية التقليدية فحسب، بل يوفر أيضًا تحسينات كبيرة في الأداء المالي. تشير هذه التقدمات إلى تحول في نظرية إدارة المخزون نحو توجيه خلق القيمة، مما يقدم نموذجًا جديدًا يدمج العمليات والمالية. عمليًا، تقدم النتائج أدوات اتخاذ قرارات قابلة للتنفيذ تمكن الشركات من تحسين كل من الكفاءة التشغيلية والصحة المالية في ظروف السوق الديناميكية.
مقدمة
في مقدمة هذه الورقة البحثية، يؤكد المؤلفون على الدور الحاسم لإدارة سلسلة التوريد، وخاصة إدارة المخزون، في تعزيز القدرة التنافسية واستدامة الشركات في سوق عالمي. يسلطون الضوء على التوازن الدقيق المطلوب للحفاظ على مستويات المخزون المثلى التي تقلل التكاليف مع تلبية الطلب المتقلب. لقد أثبتت الأساليب التقليدية لإدارة المخزون، التي تعتمد غالبًا على بيانات تاريخية ثابتة، عدم كفاءتها في مواجهة زيادة تقلبات السوق والأحداث العالمية غير المتوقعة، مما يؤدي إلى مشاكل مثل تراكم المخزون والنقص. يشير المؤلفون إلى أنه على الرغم من تطور النظريات الأساسية في إدارة المخزون، فإن تعقيد سلاسل التوريد الحديثة قد تجاوز قدرات الأساليب التحليلية التقليدية، مما يستلزم استكشاف تقنيات المحاكاة والتحسين.
تقدم الورقة التعلم العميق المعزز (DRL) كتحسين واعد في معالجة تحديات إدارة المخزون تحت عدم اليقين. ينتقد المؤلفون التطبيقات الحالية لـ DRL لتركيزها الضيق على تكاليف التشغيل، مما يتجاهل التأثير الكبير لقرارات المخزون على السيولة المالية ودورة تحويل النقد (CCC). يجادلون من أجل إطار موحد يدمج المقاييس المالية في آليات تعلم DRL، مما يمكّن من نهج أكثر شمولاً في اتخاذ قرارات المخزون. تقترح الدراسة إطار تحسين تعاوني يجمع بين محاكاة الأحداث المنفصلة وDRL، بهدف تطوير استراتيجيات مخزون تستجيب لتقلبات الطلب وتعزز أيضًا إدارة التدفق النقدي. للتحقق من نهجهم، يخطط المؤلفون لاستخدام مجموعة بيانات دقة التنبؤ Walmart M5، التي تقدم بيانات مبيعات واسعة للاختبار التجريبي في بيئة محاكاة واقعية.
طرق
تناقش هذه القسم تطبيق طرق تحسين المحاكاة في إدارة المخزون، مع تسليط الضوء على مزاياها مقارنة بالنماذج التحليلية التقليدية. تستخدم هذه الطرق محاكاة الأحداث المنفصلة لنمذجة ديناميكيات سلسلة التوريد المعقدة بدقة، مع دمج العمليات العشوائية المتعلقة بالطلبات والمخزون وأحداث الطلب. ومع ذلك، تفتقر نماذج المحاكاة إلى قدرات التحسين المدمجة ويجب دمجها مع خوارزميات ميتاheuristic، مثل البحث المحظور والخوارزميات الجينية، لتعزيز فعاليتها. على الرغم من إمكانياتها، تتعرض تحسينات المحاكاة للنقد لكونها “غير متصلة” بطبيعتها، مما يتطلب موارد حاسوبية واسعة ووقتًا للتحسين، مما يحد من القدرة على التكيف في ظروف السوق المتغيرة بسرعة.
تحدد قسم المنهجية دراسة تجريبية باستخدام مجموعة بيانات دقة التنبؤ M5 للتحقق من إطار العمل المقترح للتعاون بين المالية والتشغيل. تقارن الدراسة بين أربع طرق أساسية، بما في ذلك استراتيجية (s, S) التقليدية وشبكة Q المتكررة العميقة (DRL-Inv)، مقابل نهج نموذج تقييم الحمل المعرفي الديناميكي (CLDAM-PPO). تشير النتائج إلى أن CLDAM-PPO يقلل بشكل كبير من التكاليف الإجمالية ويحسن وقت دورة تحويل النقد (CCC) مع الحفاظ على مستويات خدمة عالية. يُعزى هذا النجاح إلى التصميم المبتكر لدالة مكافأة تآزرية توازن بين الكفاءة التشغيلية والصحة المالية، مما يفعّل الادعاءات النظرية بشأن كفاءة رأس المال العامل. تساهم الدراسة في دمج إدارة العمليات والمالية الشركات، موسعة حدود التعلم العميق المعزز (DRL) في إدارة سلسلة التوريد وتقديم إطار لاتخاذ القرارات التكيفية في بيئات غير مؤكدة. تشمل اتجاهات البحث المستقبلية تعزيز قدرة النموذج على التكيف مع المعلمات المالية الديناميكية واستكشاف قابليته للتطبيق عبر صناعات مختلفة.
مناقشة
تسلط قسم المناقشة في الورقة الضوء على التقدمات والقيود في التعلم المعزز (RL)، وخاصة التعلم العميق المعزز (DRL)، في تحسين المخزون. يؤكد على أنه بينما أحدث DRL ثورة في النهج المتبع في مشاكل اتخاذ القرار المتسلسل، تركز العديد من الدراسات الحالية، مثل تلك التي أجراها Oroojlooyjadid وآخرون (2022) وSelukar وآخرون (2022)، بشكل أساسي على تحسين تكاليف التشغيل دون معالجة الآثار المالية لقرارات المخزون بشكل كافٍ. يجادل المؤلفون بأن دوال المكافأة التقليدية، التي تأخذ عادةً في الاعتبار فقط التركيبات الخطية لتكاليف التشغيل، تتجاهل التأثير الأوسع على سيولة الشركة المالية، مما يحد من فعالية استراتيجيات DRL في تحقيق توازن بين الكفاءة التشغيلية والصحة المالية.
لمعالجة هذه الفجوة، تقدم الدراسة إطار عمل جديد، CLDAM-PPO، الذي يدمج الأهداف التشغيلية والمالية من خلال دالة مكافأة تآزرية. تتضمن هذه الدالة مقاييس مالية رئيسية، مثل دورة تحويل النقد (CCC)، جنبًا إلى جنب مع تكاليف التشغيل التقليدية. يتم التحقق من صحة الإطار باستخدام بيئة محاكاة عالية الدقة تعتمد على مجموعة بيانات M5، مما يظهر تحسينات كبيرة في التكاليف الإجمالية المرتبطة ومقاييس التدفق النقدي مقارنة باستراتيجيات المخزون التقليدية وطرق DRL الحالية. تشير النتائج إلى أن إطار CLDAM-PPO لا يعزز الأداء التشغيلي فحسب، بل يحسن أيضًا النتائج المالية، مما يوفر حلاً شاملاً لتحديات إدارة المخزون في بيئات البيع بالتجزئة الديناميكية.
القيود
تنشأ قيود نظرية التحكم في المخزون التقليدية من اعتمادها الأساسي على افتراضات صارمة لا تنطبق في السيناريوهات المعقدة الواقعية. قدمت النماذج المبكرة، مثل نموذج كمية الطلب الاقتصادية (EOQ) واستراتيجية (s, S)، أطرًا لتقليل التكاليف وإدارة المخزون تحت ظروف الطلب المستقرة. ومع ذلك، تكافح هذه النماذج، كما أبرزها Mauldin (2017)، مع قابلية التوسع عندما تواجه أنظمة متعددة المنتجات ومتعددة المستويات أو طلب غير ثابت، مما يؤدي إلى تحديات حسابية تعرف باسم “الكارثة البُعدية”.
علاوة على ذلك، بينما كانت الاستراتيجيات التي طورها Silver وآخرون (1998) تهدف إلى معالجة هذه التعقيدات، أشار Graves (1999) إلى أن فعاليتها غالبًا ما تتعرض للخطر بسبب الاعتماد على بيانات تاريخية ثابتة، خاصة في البيئات الديناميكية التي تتميز بالترويج أو تقلبات السوق. انتقد Armenzoni وآخرون (2015) الاستراتيجيات التقليدية لمحدوديتها، مما يؤدي إلى تحيزات منهجية وعدم الاستفادة الكافية من المعلومات في الوقت الحقيقي. وبالتالي، لا يزال هناك فجوة كبيرة بين الصرامة التحليلية للنظريات التقليدية في إدارة المخزون وقابليتها العملية في سلاسل التوريد المعقدة اليوم.
DOI: https://doi.org/10.1504/ijspm.2026.152090
Publication Date: 2026-01-01
Author(s): Zhifeng Qu
Primary Topic: Supply Chain and Inventory Management
Overview
This research addresses the disconnect between operational and financial objectives in supply chain inventory management by introducing a novel framework that combines discrete-event simulation with deep reinforcement learning (DRL). The proposed model, termed the cognitive load dynamic assessment model-proximal policy optimization (CLDAM-PPO), utilizes a dual-objective reward function that incorporates traditional inventory costs (holding, shortage, ordering) alongside the financial metric of the cash conversion cycle (CCC). Evaluated using the M5 forecasting accuracy dataset, the model achieved a total cost of $285.4 \pm 8.7$ (in thousands), significantly outperforming state-of-the-art baselines (p < 0.01) while maintaining a service level of 98.2% and reducing the CCC to 35.2 days. The study's conclusions emphasize the framework's ability to bridge the operational-financial decision-making gap in inventory management. By integrating operational cost considerations with financial liquidity, the CLDAM-PPO framework not only enhances traditional operational metrics but also provides substantial improvements in financial performance. This advancement signifies a shift in inventory management theory towards a value creation orientation, offering a new paradigm that merges operations and finance. Practically, the findings present actionable decision-making tools that empower enterprises to improve both operational efficiency and financial health in dynamic market conditions.
Introduction
In the introduction of this research paper, the authors emphasize the critical role of supply chain management, particularly inventory management, in enhancing the competitiveness and sustainability of enterprises in a globalized market. They highlight the delicate balance required to maintain optimal inventory levels that minimize costs while meeting fluctuating demand. Traditional inventory management approaches, which often rely on static historical data, have proven inadequate in the face of increasing market volatility and unexpected global events, leading to issues such as inventory backlogs and shortages. The authors note that while foundational theories in inventory management have evolved, the complexity of modern supply chains has outstripped the capabilities of conventional analytical methods, necessitating the exploration of simulation and optimization techniques.
The paper introduces deep reinforcement learning (DRL) as a promising advancement in addressing the challenges of inventory management under uncertainty. The authors critique existing DRL applications for their narrow focus on operational costs, which neglects the significant impact of inventory decisions on financial liquidity and the cash conversion cycle (CCC). They argue for a unified framework that integrates financial metrics into DRL learning mechanisms, enabling a more comprehensive approach to inventory decision-making. The study proposes a collaborative optimization framework that combines discrete-event simulation with DRL, aiming to develop inventory strategies that not only respond to demand fluctuations but also enhance cash flow management. To validate their approach, the authors plan to utilize the Walmart M5 forecasting accuracy dataset, which offers extensive sales data for empirical testing in a realistic simulation environment.
Methods
The section discusses the application of simulation optimization methods in inventory management, highlighting their advantages over traditional analytical models. These methods utilize discrete-event simulation to accurately model complex supply chain dynamics, incorporating stochastic processes related to orders, inventory, and demand events. However, simulation models lack inherent optimization capabilities and must be combined with metaheuristic algorithms, such as taboo search and genetic algorithms, to enhance their effectiveness. Despite their potential, simulation optimization is criticized for being inherently “offline,” requiring extensive computational resources and time for optimization, which limits adaptability in rapidly changing market conditions.
The methodology section outlines an empirical study using the M5 forecasting-accuracy dataset to validate a proposed financial-operational co-optimization framework. The study compares four baseline methods, including a classical (s, S) strategy and a deep recurrent Q-network (DRL-Inv), against the novel cognitive load dynamic assessment MModel-PPO (CLDAM-PPO) approach. The results indicate that CLDAM-PPO significantly reduces total costs and improves cash conversion cycle (CCC) time while maintaining high service levels. This success is attributed to the innovative design of a synergistic reward function that balances operational efficiency with financial health, operationalizing the theoretical claims regarding working capital efficiency. The study contributes to the integration of operations management and corporate finance, extending the boundaries of deep reinforcement learning (DRL) in supply chain management and providing a framework for adaptive decision-making in uncertain environments. Future research directions include enhancing the model’s adaptability to dynamic financial parameters and exploring its applicability across various industries.
Discussion
The discussion section of the paper highlights the advancements and limitations of reinforcement learning (RL), particularly deep reinforcement learning (DRL), in inventory optimization. It emphasizes that while DRL has revolutionized the approach to sequential decision-making problems, many existing studies, such as those by Oroojlooyjadid et al. (2022) and Selukar et al. (2022), primarily focus on operational cost optimization without adequately addressing the financial implications of inventory decisions. The authors argue that traditional reward functions, which typically consider only linear combinations of operational costs, neglect the broader impact on a firm’s financial liquidity, thereby limiting the effectiveness of DRL strategies in achieving a balance between operational efficiency and financial health.
To address this gap, the study introduces a novel framework, CLDAM-PPO, which integrates operational and financial objectives through a synergistic reward function. This function incorporates key financial metrics, such as the cash conversion cycle (CCC), alongside traditional operational costs. The framework is validated using a high-fidelity simulation environment based on the M5 dataset, demonstrating significant improvements in total associated costs and cash flow metrics compared to classical inventory strategies and existing DRL methods. The findings suggest that the CLDAM-PPO framework not only enhances operational performance but also optimizes financial outcomes, thereby providing a comprehensive solution to the challenges of inventory management in dynamic retail environments.
Limitations
The limitations of classical inventory control theory stem from its foundational reliance on strict assumptions that do not hold in complex, real-world scenarios. Early models, such as the economic order quantity (EOQ) model and the (s, S) strategy, provided frameworks for minimizing costs and managing inventory under stable demand conditions. However, these models, as highlighted by Mauldin (2017), struggle with scalability when faced with multi-product, multi-tier systems or non-stationary demand, leading to computational challenges known as ‘dimensional catastrophe.’
Moreover, while heuristics developed by Silver et al. (1998) aimed to address these complexities, Graves (1999) noted that their effectiveness is often compromised by reliance on static historical data, particularly in dynamic environments characterized by promotions or market fluctuations. Armenzoni et al. (2015) further critiqued traditional strategies for their rigidity, which results in systematic biases and underutilization of real-time information. Consequently, a significant gap remains between the analytical rigor of classical inventory theories and their practical applicability in today’s intricate supply chains.
