المعالجة غير الخطية باستخدام البصريات الخطية
Nonlinear processing with linear optics

شارك:
المجلة: Nature Photonics، المجلد: 18، العدد: 10
DOI: https://doi.org/10.1038/s41566-024-01494-z
PMID: https://pubmed.ncbi.nlm.nih.gov/39372106
تاريخ النشر: 2024-07-31
المؤلف: Mustafa Yildirim وآخرون
الموضوع الرئيسي: الشبكات العصبية وحوسبة الخزانات

طرق

في هذه الدراسة، تضمنت الإعدادات التجريبية ليزر Solstis M2 يعمل على طول موجي قدره $\lambda = 850 \, \text{nm}$. كانت التكوينات البصرية تشمل مرآة بعرض 11 مم، مما يسهل أربعة انعكاسات لشعاع الليزر. تم استخدام جهاز المرايا الرقمية (DMD) لنقل الشعاع إلى وحدة تعديل الضوء المكاني (SLM) باستخدام تصوير 4f، مما يسمح بتشكيل الشعاع بشكل مرن إلى مربع بطول جانب 2.4 مم، مما يتوافق مع منطقة 300 × 300 بكسل على SLM، الذي لديه مسافة بكسل قدرها $\Lambda = 8 \, \mu\text{m}$. تم وضع المرآة على بعد 15.2 مم من SLM لضمان أن الانكسار من بكسل الزاوية الواحدة يمكن أن يصل إلى الزاوية المقابلة من الطبقة التالية، مما يفي بشرط $d \times \frac{\lambda}{\Lambda} \geq 2.4 \, \text{mm}$.

استخدم الإعداد انعكاس النظام الصفري من SLM، الذي كان لديه قدرة انعكاسية قدرها 69%، مما أدى إلى كفاءة نقل إجمالية قدرها 23% بعد أربع ارتدادات. تم تكبير الشعاع الناتج بمعامل 1.2 وتم التقاطه بواسطة كاميرا CMOS مع مسافة بكسل قدرها 3.45 ميكرومتر، تغطي منطقة 834 × 834 بكسل. شمل اكتساب الصورة استخدام تجميع متوسط لتغيير حجم الصور إلى أبعاد 4 × 4 أو 8 × 8 بكسل. بالنسبة لمجموعة بيانات Imagenette، تم استخدام عينات التدريب والاختبار الكاملة، بينما لمجموعات بيانات Digit وFashion MNIST، تم استخدام مجموعة التدريب الكاملة (60,000 عينة) للمحاكاة، مع تخصيص أول 10,000 عينة لإعادة تدريب الأوزان الرقمية و2,500 عينة من مجموعة الاختبار للاختبار الأعمى للنتائج التجريبية.

نتائج

في هذا القسم، يستكشف المؤلفون تأثيرات تكوين الطبقات وتخصيص البيانات على أداء إطار عمل nPOLO في مهام التصنيف البصري. من خلال الحفاظ على عدد ثابت من الطبقات والبكسلات مع تغيير تخصيص البيانات – باستخدام الطبقة الأولى فقط للبيانات والطبقات اللاحقة لمتغيرات التحيز القابلة للتدريب – أجروا تجارب على مجموعات بيانات Imagenette وFashion MNIST وDigit MNIST. تشير النتائج، الموضحة في الشكل 2، إلى أن تكرار البيانات عبر طبقات متعددة يعزز دقة التصنيف، مع زيادة عدد الطبقات التي تساهم بشكل أكبر في تحسين الأداء. على العكس من ذلك، أظهرت التكوينات التي لا تحتوي على تكرار البيانات انخفاضًا في الدقة عندما تم الاحتفاظ بعدد الطبقات ثابتًا.

استخدم المؤلفون طريقة انتشار الشعاع (BPM) لمحاكاة انتشار الضوء، مما يسمح بتحسين المتغيرات القابلة للتدريب عبر الانحدار العشوائي. لاحظوا أنه بينما كانت النتائج التجريبية تؤكد باستمرار إطار عمل nPOLO، كانت الدقة أقل من تلك التي تم الحصول عليها من خلال المحاكاة، خاصة في السيناريوهات التي تتضمن تكرار البيانات. يُعزى هذا التباين إلى العيوب في التنفيذ الفيزيائي، مثل الاستجابات الطورية غير المثالية لوحدة تعديل الضوء المكاني (SLM). كما قارن المؤلفون أداء إطار عمل nPOLO ضد شبكة عصبية تلافيفية بسيطة، ووجدوا نتائج قابلة للمقارنة، والتي تم تفصيلها في القسم التكميلي 4. ستستكشف الدراسات المستقبلية الجوانب الهندسية لإطار عمل nPOLO لمعالجة هذه التباينات.

مناقشة

في هذا القسم، يناقش المؤلفون أداء وموثوقية معالجة غير الخطية مع البصريات الخطية فقط (nPOLO) لإجراء الشبكات العصبية البصرية. تكشف الدراسة أنه عندما يكون لدى كلا النظامين عدد متساوٍ من درجات الحرية، فإن استخدام استراتيجية تكرار البيانات يؤدي باستمرار إلى دقة أعلى ومرونة أكبر تجاه العيوب التجريبية والضوضاء. يدمج إطار عمل nPOLO عدة مستويات بيانات باستخدام وحدة تعديل الضوء السائل البلوري (SLM) ومرآة، مما يسهل تعديل الضوء في وقت واحد ويمكّن من التحولات الخطية وغير الخطية للبيانات. تشير النتائج إلى أن زيادة عدد طبقات التعديل وتكرار البيانات تعزز توليد غير الخطيات من الدرجة العليا، والتي تعتبر حاسمة لمهام التصنيف المعقدة.

كما يبرز المؤلفون أهمية متغيرات التوسع والتحيز القابلة للتدريب لقيم البكسل المعروضة على SLM، مما يسمح بتحولات قابلة للبرمجة. تُظهر دراستهم للتوسع علاقة قانون القوة بين حجم النموذج والأداء، مشابهة للنتائج في الشبكات العصبية الرقمية. من الجدير بالذكر أن إدخال تكرار البيانات لا يحسن الدقة فحسب، بل يعزز أيضًا الموثوقية ضد الضوضاء، حيث أظهرت التكوينات ذات البيانات المكررة حساسية أقل للتغيرات التجريبية. تشير النتائج إلى أن nPOLO يمكن أن تولد غير خطية بصرية بفعالية دون مكونات إلكترونية، مما يجعلها نهجًا واعدًا لتطوير الشبكات العصبية البصرية. قد تستكشف الأبحاث المستقبلية تنفيذات بديلة للمتغيرات القابلة للتدريب لتعزيز قدرات الإطار بشكل أكبر.

Journal: Nature Photonics, Volume: 18, Issue: 10
DOI: https://doi.org/10.1038/s41566-024-01494-z
PMID: https://pubmed.ncbi.nlm.nih.gov/39372106
Publication Date: 2024-07-31
Author(s): Mustafa Yildirim et al.
Primary Topic: Neural Networks and Reservoir Computing

Methods

In this study, the experimental setup involved a continuous-wave Solstis M2 laser operating at a wavelength of $\lambda = 850 \, \text{nm}$. The optical configuration included a mirror with a width of 11 mm, facilitating four reflections of the laser beam. A digital micromirror device (DMD) was employed to relay the beam to a spatial light modulator (SLM) using 4f imaging, allowing for flexible beam shaping into a square of side length 2.4 mm, corresponding to a 300 × 300-pixel area on the SLM, which has a pixel pitch of $\Lambda = 8 \, \mu\text{m}$. The mirror was positioned 15.2 mm from the SLM to ensure that diffraction from one corner pixel could reach the opposite corner of the subsequent layer, satisfying the condition $d \times \frac{\lambda}{\Lambda} \geq 2.4 \, \text{mm}$.

The setup utilized zero-order reflection from the SLM, which had a reflectivity of 69%, resulting in a total transmission efficiency of 23% after four bounces. The output beam was magnified by a factor of 1.2 and captured by a complementary metal-oxide-semiconductor (CMOS) camera with a pixel pitch of 3.45 μm, covering an area of 834 × 834 pixels. Image acquisition involved average pooling to resize images to dimensions of either 4 × 4 or 8 × 8 pixels. For the Imagenette dataset, the complete training and test samples were utilized, while for the Digit and Fashion MNIST datasets, the entire training set (60,000 samples) was used for simulations, with the first 10,000 samples reserved for re-training digital weights and 2,500 samples from the test set for blind testing of experimental results.

Results

In this section, the authors investigate the effects of layer configuration and data allocation on the performance of the nPOLO framework in optical classification tasks. By maintaining a constant number of layers and pixels while varying data allocation—specifically using only the first layer for data and subsequent layers for trainable bias parameters—they conducted experiments on the Imagenette, Fashion MNIST, and Digit MNIST datasets. The results, illustrated in Figure 2, indicate that repeating data across multiple layers enhances classification accuracy, with increased layer counts further contributing to improved performance. Conversely, configurations without data repetition showed a decline in accuracy when the number of layers was held constant.

The authors employed a beam propagation method (BPM) for simulating light propagation, allowing for the optimization of trainable parameters via stochastic gradient descent. They noted that while experimental results consistently validated the nPOLO framework, the accuracies were lower than those obtained through simulations, particularly in scenarios involving data repetition. This discrepancy is attributed to imperfections in the physical implementation, such as non-ideal phase responses of the spatial light modulator (SLM). The authors also compared the nPOLO framework’s performance against a simple convolutional neural network, finding comparable results, which are detailed in Supplementary Section 4. Future studies will further explore the engineering aspects of the nPOLO framework to address these discrepancies.

Discussion

In this section, the authors discuss the performance and robustness of the nonlinear processing with only linear optics (nPOLO) framework for optical neural networks. The study reveals that when both systems have an equal number of degrees of freedom, employing a data repetition strategy consistently yields higher accuracy and greater resilience to experimental imperfections and noise. The nPOLO framework integrates multiple data planes using a liquid-crystal spatial light modulator (SLM) and a mirror, facilitating simultaneous modulation of light and enabling both linear and nonlinear transformations of data. The results indicate that increasing the number of modulation layers and data repetitions enhances the generation of higher-order nonlinearities, which are crucial for complex classification tasks.

The authors also highlight the significance of trainable scaling and bias parameters for pixel values displayed on the SLM, which allow for programmable transformations. Their scaling study demonstrates a power-law relationship between model size and performance, akin to findings in digital neural networks. Notably, the introduction of data repetition not only improves accuracy but also enhances robustness against noise, as configurations with repeated data showed less sensitivity to experimental variations. The findings suggest that nPOLO can effectively generate optical nonlinearity without electronic components, making it a promising approach for advancing optical neural networks. Future research may explore alternative implementations of trainable parameters to further enhance the framework’s capabilities.

شارك: