معالجة المعلومات البصرية متعددة الأطوال الموجية باستخدام التعلم العميق المعزز
Multi-wavelength optical information processing with deep reinforcement learning

شارك:
المجلة: Light Science & Applications، المجلد: 14، العدد: 1
DOI: https://doi.org/10.1038/s41377-025-01846-6
PMID: https://pubmed.ncbi.nlm.nih.gov/40229251
تاريخ النشر: 2025-04-15
المؤلف: Qiuquan Yan وآخرون
الموضوع الرئيسي: الشبكات العصبية وحوسبة الخزانات

نظرة عامة

يقدم هذا القسم نهجًا جديدًا لتعزيز أنظمة معالجة المعلومات البصرية متعددة الأطوال الموجية، والتي تعتبر جزءًا لا يتجزأ من الشبكات العصبية البصرية ومعالجة الإشارات العريضة النطاق. غالبًا ما تعيق فعالية هذه الأنظمة الاستجابات الانتقائية للتردد الناتجة عن عوامل مختلفة، بما في ذلك ظروف التصنيع والبيئة. لمعالجة هذه التحديات، يقدم البحث طريقة معايرة تعتمد على التعلم المعزز العميق (DRC)، مستلهمًا من استراتيجية تدريب تدرج السياسة الحتمي العميق. تتعلم هذه الطريقة بشكل مستقل من النظام، مما يجمع المعرفة التجريبية التي تؤدي إلى تحسين استراتيجيات المعايرة والقدرة على التكيف مقارنةً بالتقنيات التقليدية.

تم تطبيق طريقة DRC بنجاح على أنظمة تستخدم ألياف تعويض التشتت، ومصفوفات الرنانات الدقيقة، ومصفوفات مقياس تداخل ماخ-زنر، محققة الوظائف اللازمة لمعالجة الإشارات في 21 تكرارًا فقط. لا تعزز هذه الطريقة كفاءة التحكم والدقة فحسب، بل تظهر أيضًا وعدًا للتطبيقات في تسريع حساب الالتفاف البصري، ومعالجة الإشارات الضوئية الميكروويف، وتوجيه الشبكات البصرية.

مقدمة

تناقش مقدمة هذه الورقة البحثية أهمية أنظمة معالجة المعلومات البصرية متعددة الأطوال الموجية، التي تعتبر محورية في تطبيقات مثل الحوسبة ومعالجة الإشارات العريضة النطاق. تستفيد هذه الأنظمة من الخصائص الفريدة للضوء وتقسيم الطول الموجي، مما يوفر مزايا مثل عرض النطاق الترددي العالي، والانخفاض في التوهين، وكفاءة الطاقة. ومع ذلك، تواجه التطبيقات العملية تحديات بسبب الاستجابات الانتقائية للتردد التي يمكن أن تؤدي إلى أخطاء حسابية، خاصة في الأنظمة التي تستخدم ألياف تعويض التشتت (DCF)، والرنانات الدقيقة (MRR)، ومقاييس تداخل ماخ-زنر (MZI). تشمل العوامل المساهمة في هذه المشكلات عيوب التصنيع، والاستجابات الترددية غير المتجانسة من وسائط الكسب، والتأثيرات البيئية.

لمعالجة هذه التحديات، تقدم الورقة طريقة معايرة جديدة تُسمى المعايرة المعتمدة على التعلم المعزز العميق (DRC). تستخدم هذه الطريقة استراتيجيات تدرج السياسة الحتمي العميق (DDPG) لتعزيز القدرة على التكيف وتقليل أخطاء المعايرة في الأنظمة متعددة الأطوال الموجية. يقوم نموذج DRC بالتدريب المستمر وتحسين استراتيجيات المعايرة الخاصة به، مما يسمح له بالاستجابة بفعالية للاضطرابات البيئية. أظهرت التحقق التجريبي عبر ثلاثة أنظمة متعددة الأطوال الموجية تحسينات كبيرة في كفاءة المعايرة والدقة، مع تقليل خطأ الإخراج بنسبة تصل إلى 85.4% مقارنةً بأساليب المعايرة التقليدية. تؤكد النتائج على إمكانية DRC في تعزيز موثوقية ودقة الأنظمة البصرية متعددة الأطوال الموجية.

الطرق

يصف قسم الطرق نموذج معالجة المعايرة والنظام متعدد الأطوال الموجية الذي تم تطويره في هذه الدراسة، كما هو موضح في الشكل 5. يتكون النظام من أربعة وحدات: حوامل ليزر متعددة الأطوال الموجية، وحدة تعديل وترميز، وحدة معالجة المعلومات، ووحدة كشف. يتم استخدام مشط تردد ضوئي (OFC) في مصفوفات الليزر ذات التغذية المرتدة الموزعة (DFB) لتوفير عدة حوامل متباعدة بالتساوي، بينما تتطلب الرنانات الدقيقة (MRRs) ومقاييس تداخل ماخ-زنر (MZIs) طاقة إخراج مستقرة وأطوال موجية محددة. تتضمن عملية المعايرة تبادل البيانات بين وحدة التعديل والترميز ووحدة الكشف، بدءًا من إصدار إشارة معيارية لتقييم إشارات الإخراج بحثًا عن الأخطاء.

تستخدم المعايرة خوارزمية المعايرة الديناميكية المكافأة (DRC)، التي تعدل بشكل تكراري معلمات المعايرة بناءً على الفروقات بين الإخراج الفعلي والإخراج المثالي. يتم هيكلة نموذج DRC حول مكونات الحالة، والإجراء، والمكافأة، حيث تمثل الحالة الإخراج العادي، ويتوافق الإجراء مع التغييرات في معلمات المعايرة، وتحدد دالة المكافأة مدى توافق الإخراجات. يستخدم النموذج إطار عمل مزدوج الشبكة مع شبكة ممثل لإنشاء استراتيجيات الإجراءات وشبكة ممثل مستهدفة لتحسين سياسات المعايرة. يخزن مخزن إعادة التشغيل بيانات المعايرة لتعزيز تدريب نموذج الممثل، مما يسهل التقارب نحو استراتيجية معايرة فعالة تقلل الأخطاء عبر أنظمة بصرية مختلفة. كما يتم مناقشة المقارنات مع خوارزميات أخرى، مثل الخوارزمية الجينية (GA)، وتدرج الاضطراب المتزامن (SPGD)، والتناسب-التكامل-المشتق (PID)، مع توفير مزيد من التفاصيل في الملاحظات التكميلية.

النتائج

يستكشف قسم النتائج في الدراسة تأثير طرق المعايرة المختلفة على دقة الأنظمة متعددة الأطوال الموجية المستخدمة في معالجة المعلومات، باستخدام نماذج الشبكات العصبية المستندة إلى هذه الأنظمة. تتضمن النماذج طبقات التلافيف التي تستفيد من نوى تلافيف مختلفة مشتقة من ثلاثة أنواع من الأنظمة متعددة الأطوال الموجية. تم تقييم أداء هذه النماذج باستخدام مجموعات بيانات MNIST وUrbanSound8K، حيث كانت الشبكة العصبية المثالية بمثابة معيار للمقارنة. تشير النتائج إلى أن النموذج المثالي حقق أعلى دقة تصنيف، يليه النموذج الذي تم معايرته باستخدام طريقة DRC، بينما أظهر النموذج المعاير باستخدام PID أدنى دقة. يشير ذلك إلى أن طريقة DRC تتفوق على نهج PID في تعزيز أداء النظام.

تم إجراء مزيد من التحليل على الأنظمة المستندة إلى مصفوفة MRR ومصفوفة MZI، حيث أظهرت طريقة DRC كفاءة معايرة أعلى بكثير مقارنةً بطريقة PID وخوارزمية GA. على وجه التحديد، أدت طريقة DRC إلى تقليل الأخطاء ودقة التصنيف أقرب إلى النموذج المثالي، مما يبرز العلاقة بين تقليل أخطاء المعايرة وتحسين الدقة في مهام معالجة المعلومات البصرية متعددة الأطوال الموجية. بشكل عام، تؤكد النتائج فعالية طريقة DRC في تعزيز موثوقية أنظمة معالجة المعلومات البصرية.

المناقشة

تقيّم قسم المناقشة في هذه الورقة البحثية كفاءة المعايرة لأربع طرق—طيف مزدوج (DRC)، خوارزمية جينية (GA)، تدرج موازٍ عشوائي (SPGD)، وتناسب-تكامل-مشتق (PID)—عبر ثلاثة أنظمة متعددة الأطوال الموجية: الأنظمة المستندة إلى DCF، والأنظمة المستندة إلى مصفوفة MRR، والأنظمة المستندة إلى مصفوفة MZI. كانت طريقة DRC، التي تدمج المعرفة السابقة وتستخدم نهج التعلم المعزز العميق، تتفوق باستمرار على الطرق الأخرى من حيث سرعة المعايرة ودقتها. على وجه الخصوص، حققت طريقة DRC تخفيضات كبيرة في أخطاء الإخراج، حيث انخفضت أخطاء نظام DCF إلى 22.0% من قيمها الأصلية، وبلغت تباين الإخراج لمصفوفة MRR 3.89 × 10⁻⁴، وتم تقليل الخطأ النسبي للإخراج لمصفوفة MZI إلى 14.6% من نتائج المعايرة اليدوية.

تسلط الدراسة الضوء على نقاط القوة والضعف لكل طريقة معايرة. بينما تعتبر الخوارزميات الحدسية مثل GA وSPGD مفيدة للمشكلات المعقدة، إلا أنها غالبًا ما تتطلب تكرارات واسعة وقد تواجه صعوبة مع الأنظمة الأكبر. في المقابل، تعتبر الطرق الحتمية مثل PID أسرع ولكن تعتمد بشكل كبير على المعرفة الخبيرة لضبط المعلمات. إن قدرة طريقة DRC على تحقيق حالات المعايرة المرغوبة بسرعة من خلال اختيار السياسات الاستراتيجية، على الرغم من الحاجة إلى التدريب المسبق، تجعلها الحل الأكثر فعالية للتطبيقات في الوقت الحقيقي في الحوسبة البصرية والضوئيات الميكروويف. بشكل عام، تؤكد النتائج على إمكانية خوارزمية DRC في تعزيز عمليات المعايرة في الأنظمة متعددة الأطوال الموجية، مما يمهد الطريق لتحسين الأداء في التطبيقات العملية.

Journal: Light Science & Applications, Volume: 14, Issue: 1
DOI: https://doi.org/10.1038/s41377-025-01846-6
PMID: https://pubmed.ncbi.nlm.nih.gov/40229251
Publication Date: 2025-04-15
Author(s): Qiuquan Yan et al.
Primary Topic: Neural Networks and Reservoir Computing

Overview

The section presents a novel approach to enhancing multi-wavelength optical information processing systems, which are integral to optical neural networks and broadband signal processing. The effectiveness of these systems is often hindered by frequency-selective responses arising from various factors, including fabrication and environmental conditions. To address these challenges, the study introduces a deep reinforcement learning calibration (DRC) method, drawing inspiration from the deep deterministic policy gradient training strategy. This method autonomously learns from the system, accumulating experiential knowledge that leads to improved calibration strategies and adaptability compared to conventional techniques.

The DRC method has been successfully applied to systems utilizing dispersion compensating fiber, micro-ring resonator arrays, and Mach-Zehnder interferometer arrays, achieving the necessary signal processing functions within just 21 iterations. This approach not only enhances control efficiency and accuracy but also shows promise for applications in optical convolution computation acceleration, microwave photonic signal processing, and optical network routing.

Introduction

The introduction of this research paper discusses the significance of multi-wavelength optical information processing systems, which are pivotal in applications such as computing and broadband signal processing. These systems leverage the unique properties of light and wavelength division multiplexing, offering advantages like high bandwidth, low attenuation, and energy efficiency. However, practical implementations face challenges due to frequency-selective responses that can lead to computational errors, particularly in systems utilizing dispersive compensating fiber (DCF), micro-ring resonators (MRR), and Mach-Zehnder interferometers (MZI). Factors contributing to these issues include fabrication defects, non-uniform frequency responses from gain media, and environmental influences.

To address these challenges, the paper introduces a novel calibration method termed Deep Reinforcement learning-based Calibration (DRC). This approach utilizes deep deterministic policy gradient (DDPG) strategies to enhance adaptability and reduce calibration errors in multi-wavelength systems. The DRC model continuously trains and optimizes its calibration strategies, allowing it to effectively respond to environmental disturbances. Experimental validation across three multi-wavelength systems demonstrated significant improvements in calibration efficiency and accuracy, with reductions in output error of up to 85.4% compared to traditional calibration methods. The findings underscore the potential of DRC to enhance the reliability and precision of multi-wavelength optical systems.

Methods

The methods section describes the calibration processing model and multi-wavelength system developed in this study, illustrated in Figure 5. The system consists of four modules: multi-wavelength laser carriers, a modulation and coding unit, an information processing unit, and a detection unit. An optical frequency comb (OFC) is utilized in the distributed feedback (DFB) laser arrays to provide multiple equally spaced carriers, while micro-ring resonators (MRRs) and Mach-Zehnder interferometers (MZIs) require stable output power and specific wavelengths. The calibration process involves data exchange between the modulation and coding unit and the detection unit, starting with the emission of a standard signal to evaluate output signals for errors.

The calibration employs a Dynamic Reward Calibration (DRC) algorithm, which iteratively adjusts calibration parameters based on discrepancies between the actual output and the ideal output. The DRC model is structured around state, action, and reward components, where the state represents the normalized output, the action corresponds to changes in calibration parameters, and the reward function quantifies the alignment of outputs. The model utilizes a dual-network framework with an actor network for generating action strategies and a target actor network for optimizing calibration policies. A replay buffer stores calibration data to enhance the training of the actor model, facilitating convergence toward an effective calibration strategy that minimizes errors across various optical systems. Comparisons with other algorithms, such as Genetic Algorithm (GA), Simultaneous Perturbation Gradient Descent (SPGD), and Proportional-Integral-Derivative (PID), are also discussed, with further details provided in supplementary notes.

Results

The results section of the study investigates the impact of various calibration methods on the accuracy of multi-wavelength systems used in information processing, utilizing neural network models based on these systems. The models incorporate convolutional layers that leverage different convolutional kernels derived from three types of multi-wavelength systems. The performance of these models was evaluated using the MNIST and UrbanSound8K datasets, with the ideal neural network serving as a benchmark for comparison. The findings indicate that the ideal model achieved the highest classification accuracy, followed by the model calibrated using the DRC method, while the PID-calibrated model exhibited the lowest accuracy. This suggests that the DRC method is superior to the PID approach in enhancing system performance.

Further analysis was conducted on the MRR array-based and MZI-based systems, where the DRC method demonstrated significantly higher calibration efficiency compared to the PID method and the GA algorithm. Specifically, the DRC method led to lower errors and classification accuracies closer to the ideal model, underscoring the correlation between reduced calibration errors and improved accuracy in multi-wavelength optical information processing tasks. Overall, the results affirm the effectiveness of the DRC method in enhancing the reliability of optical information processing systems.

Discussion

The discussion section of this research paper evaluates the calibration efficiency of four methods—Dual-Comb Spectroscopy (DRC), Genetic Algorithm (GA), Stochastic Parallel Gradient Descent (SPGD), and Proportional-Integral-Derivative (PID)—across three multi-wavelength systems: DCF-based, MRR array-based, and MZI array-based systems. The DRC method, which integrates prior knowledge and employs a deep reinforcement learning approach, consistently outperformed the other methods in terms of calibration speed and accuracy. Specifically, the DRC method achieved significant reductions in output errors, with the DCF system’s errors decreasing to 22.0% of their original values, the MRR array’s output variance reaching 3.89 × 10⁻⁴, and the MZI array’s relative output error minimized to 14.6% of manual calibration results.

The study highlights the strengths and weaknesses of each calibration method. While heuristic algorithms like GA and SPGD are beneficial for complex problems, they often require extensive iterations and may struggle with larger systems. In contrast, deterministic methods like PID are faster but depend heavily on expert knowledge for parameter tuning. The DRC method’s ability to rapidly achieve desired calibration states through strategic policy selection, despite needing pre-training, positions it as the most effective solution for real-time applications in optical computing and microwave photonics. Overall, the findings underscore the potential of the DRC algorithm to enhance calibration processes in multi-wavelength systems, paving the way for improved performance in practical applications.

شارك: