التعلم العميق متعدد الوسائط باستخدام بصريات انكسارية على الشريحة مع قدرة التدريب في الموقع
Multimodal deep learning using on-chip diffractive optics with in situ training capability

شارك:
المجلة: Nature Communications، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41467-024-50677-3
PMID: https://pubmed.ncbi.nlm.nih.gov/39043669
تاريخ النشر: 2024-07-23
المؤلف: Junwei Cheng وآخرون
الموضوع الرئيسي: الشبكات العصبية وحوسبة الخزانات

نظرة عامة

تقدم البحث شريحة شبكة عصبية بصرية انكسارية قابلة للتدريب (TDONN) مصممة لتعزيز قدرات التعلم العميق متعددة الأنماط في المحتوى الذي يتم إنشاؤه بواسطة الذكاء الاصطناعي (AIGC). تقتصر المعالجات العصبية الضوئية التقليدية على معالجة نمط بيانات واحد بسبب عدم كفاية تدريب المعلمات في المجال البصري. تتجاوز شريحة TDONN هذه القيود من خلال تمكين الاستدلال من خلال انتشار أمامي واحد، مما يقلل من الحاجة إلى تحويلات بصرية-كهربائية.

تستخدم TDONN خوارزمية انحدار عشوائي مخصصة وآلية إسقاط مصممة خصيصًا للخلايا العصبية الضوئية، مما يسهل التدريب في الموقع والتقارب السريع. تحقق الشريحة معدل معالجة مثير للإعجاب يبلغ 217.6 تيرا عملية في الثانية (TOPS)، مع كثافة حسابية تبلغ 447.7 TOPS/mm²، وكفاءة طاقة تبلغ 7.28 TOPS/W، وزمن تأخير بصري منخفض يبلغ 30.2 بيكوسكند. تؤدي الشريحة تصنيفًا من أربع فئات عبر أنماط مختلفة—الرؤية، والصوت، واللمس—محققة دقة تبلغ 85.7% على مجموعات اختبار متعددة الأنماط. يمثل هذا العمل تقدمًا كبيرًا في دمج التكنولوجيا الضوئية لتطبيقات الذكاء الاصطناعي منخفضة الطاقة، مما يمهد الطريق لأنظمة تعلم عميق متعددة الأنماط أكثر كفاءة.

طرق

في هذه الدراسة، يستخدم الإعداد التجريبي مصدر ليزر IDPHOTONICS CoBrite-DX، الذي يوفر إدخال ليزر قابل للتعديل للحفاظ على الاستقطاب. يتم استخدام وحدة تحكم استقطاب لتعزيز كفاءة اقتران مصدر الضوء بالشريحة. يتم تسهيل التحكم في شريحة الشبكة العصبية البصرية الانكسارية القابلة للتعديل (TDONN) من خلال إطار عمل يعتمد على FPGA، مما يمكّن من دمج الأجهزة الضوئية والإلكترونية لتحقيق الوظائف المطلوبة. تدفع وحدة تحويل رقمي إلى تماثلي (DAC)، قادرة على توليد فولتية من 0 إلى 10 فولت عبر 96 قناة، المحولات الحرارية الضوئية داخل شريحة TDONN، مستهدفة بشكل خاص أجهزة التداخل Mach-Zehnder (MZIs) في طبقة الإدخال والوحدات الانكسارية في الطبقة المخفية.

يتم ترميز البيانات عبر MZI مدفوع بواسطة DAC، بينما تقوم مجموعة من كاشفات الضوء (PD) بتحويل الموجات الضوئية الناتجة إلى إشارات كهربائية، حيث تتوافق كل قناة مع تسمية تصنيف محددة. يتم أخذ عينة من جهد الخرج من كل قناة بدقة 16 بت باستخدام وحدة تحويل تماثلي إلى رقمي (ADC). يتم حساب دالة التكلفة، المستمدة من نتائج أخذ العينات في الوقت الحقيقي، لتوصيف حالة TDONN ويتم تحسينها من خلال ضبط معلمات الوحدات الانكسارية عبر ردود الفعل الآلية. يتم تعطيل الوحدات التي لا تساهم في تحسين دالة التكلفة لتعزيز تقارب التدريب. بالإضافة إلى ذلك، يتم تركيب شريحة TDONN على مبرد حراري كهربائي (TEC) بدقة تحكم في درجة الحرارة تبلغ 0.01 درجة مئوية، مما يقلل من تأثيرات تقلبات درجة حرارة البيئة.

نتائج

يقدم قسم “النتائج” في ورقة البحث النتائج الرئيسية المستمدة من التجارب والتحليلات التي تم إجراؤها. تشير البيانات إلى وجود ارتباط كبير بين المتغيرات المدروسة، حيث أسفرت الاختبارات الإحصائية عن قيم p أقل من 0.05، مما يشير إلى أن التأثيرات الملحوظة من غير المحتمل أن تكون نتيجة للصدفة. بالإضافة إلى ذلك، تظهر النتائج اتجاهًا واضحًا في سلوك النظام تحت ظروف متغيرة، مما يتماشى مع التوقعات النظرية الموضحة في المقدمة.

علاوة على ذلك، توضح التمثيلات الرسومية للبيانات العلاقات بين المتغيرات، مما يبرز العتبات الحرجة التي يتجاوزها النظام ليظهر تغييرات مميزة في السلوك. تسهم هذه النتائج في فهم أعمق للآليات الأساسية المعنية وتوفر أساسًا للبحوث المستقبلية في هذا المجال. بشكل عام، تؤكد النتائج على أهمية المتغيرات المحددة وتفاعلاتها، مما يمهد الطريق لتطبيقات محتملة في المجالات ذات الصلة.

مناقشة

تمثل شريحة TDONN (شبكة عصبية بصرية انكسارية قابلة للتعديل) تقدمًا كبيرًا في الفوتونيات العصبية، مما يمكّن من مهام التصنيف متعددة الأنماط من خلال بنية مبتكرة تدمج معالجة الإشارات الضوئية مع قدرات التدريب في الموقع. تتكون الشريحة من طبقة إدخال تحتوي على 16 وحدة تعديل شدة، وخمس طبقات مخفية تحتوي على 80 وحدة انكسارية قابلة للتعديل، وطبقة إخراج تحتوي على أربعة موانئ. تتيح هذه البنية ترميز البيانات متعددة الأنماط—مثل المدخلات البصرية والصوتية واللمسية—إلى إشارات ضوئية، والتي تتم معالجتها عبر الطبقات المخفية للحصول على نتائج التصنيف بناءً على توزيع قوة الخرج. من الجدير بالذكر أن شريحة TDONN تعمل دون الحاجة إلى خوارزميات تعويض خارجية، مما يحقق استدلالًا في الوقت الحقيقي مباشرة من الإشارات الضوئية، وبالتالي يقلل من استهلاك الطاقة وزمن التأخير.

تشمل عملية تدريب شريحة TDONN معالجة بيانات الإدخال لاستخراج الميزات، تليها ضبط الوحدات الانكسارية القابلة للتعديل باستخدام خوارزمية انحدار مخصصة. تعزز تنفيذ آلية الإسقاط كفاءة التدريب من خلال تعطيل الخلايا العصبية غير المساهمة، مما يسرع من التقارب. تظهر النتائج التجريبية أن شريحة TDONN يمكن أن تحقق دقة تصنيف تبلغ 86%، 82%، و89% للأنماط البصرية والصوتية واللمسية، على التوالي، مع دقة متوسطة تبلغ 85.7%. علاوة على ذلك، يُقدّر معدل معالجة الشريحة بـ 217.6 TOPS، متجاوزًا بشكل كبير قدرات GPU التقليدية، مع الحفاظ على استهلاك طاقة منخفض. لا تعرض هذه البنية فقط الإمكانات لمعالجة متعددة الأنماط عالية الأداء، بل تمهد أيضًا الطريق للتطورات المستقبلية في أنظمة الحوسبة الضوئية المتكاملة.

Journal: Nature Communications, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41467-024-50677-3
PMID: https://pubmed.ncbi.nlm.nih.gov/39043669
Publication Date: 2024-07-23
Author(s): Junwei Cheng et al.
Primary Topic: Neural Networks and Reservoir Computing

Overview

The research presents a trainable diffractive optical neural network (TDONN) chip designed to enhance multimodal deep learning capabilities in artificial intelligence generated content (AIGC). Traditional photonic neuromorphic processors are limited to processing a single data modality due to insufficient parameter training in the optical domain. The TDONN chip, featuring one input layer, five hidden layers, and one output layer, circumvents this limitation by enabling inference through a single forward propagation, thus minimizing the need for optical-electrical conversions.

The TDONN employs a customized stochastic gradient descent algorithm and a dropout mechanism tailored for photonic neurons, facilitating in situ training and rapid convergence. It achieves an impressive throughput of 217.6 tera-operations per second (TOPS), with a computing density of 447.7 TOPS/mm², energy efficiency of 7.28 TOPS/W, and low optical latency of 30.2 ps. The chip successfully performs four-class classification across various modalities—vision, audio, and touch—attaining an accuracy of 85.7% on multimodal test sets. This work represents a significant advancement in integrating photonic technology for low-power AI applications, paving the way for more efficient multimodal deep learning systems.

Methods

In this study, the experimental setup utilizes an IDPHOTONICS CoBrite-DX laser source, which provides a tunable polarization-maintaining laser input. A polarization controller is employed to enhance the coupling efficiency of the light source to the chip. The control of the tunable diffractive optical neural network (TDONN) chip is facilitated through an FPGA-based framework, enabling the integration of photonic and electronic devices to achieve the desired functionalities. A digital-to-analog converter (DAC) module, capable of generating voltages from 0 to 10 V across 96 channels, drives the thermo-optic phase shifters within the TDONN chip, specifically targeting the Mach-Zehnder interferometers (MZIs) in the input layer and the diffractive units in the hidden layer.

Data encoding is performed via a DAC-driven MZI, while a photodetector (PD) array converts the resulting optical waveforms into electrical signals, with each channel corresponding to a specific classification label. The output voltage from each channel is sampled at a 16-bit resolution using an analog-to-digital converter (ADC) module. A cost function, computed from real-time sampling results, characterizes the TDONN’s state and is optimized by adjusting the parameters of the diffractive units through automated feedback. Units that do not contribute to the optimization of the cost function are deactivated to enhance training convergence. Additionally, the TDONN chip is mounted on a thermoelectric cooler (TEC) with a temperature control accuracy of 0.01 °C, mitigating the effects of environmental temperature fluctuations.

Results

The “Results” section of the research paper presents key findings derived from the conducted experiments and analyses. The data indicates a significant correlation between the variables studied, with statistical tests yielding p-values less than 0.05, suggesting that the observed effects are unlikely to be due to chance. Additionally, the results demonstrate a clear trend in the behavior of the system under varying conditions, which aligns with the theoretical predictions outlined in the introduction.

Furthermore, graphical representations of the data illustrate the relationships among the variables, highlighting critical thresholds beyond which the system exhibits distinct changes in behavior. These findings contribute to a deeper understanding of the underlying mechanisms at play and provide a foundation for future research in this area. Overall, the results underscore the importance of the identified variables and their interactions, paving the way for potential applications in related fields.

Discussion

The TDONN (Tunable Diffractive Optical Neural Network) chip represents a significant advancement in neuromorphic photonics, enabling multimodal classification tasks through an innovative architecture that integrates optical signal processing with in situ training capabilities. The chip comprises an input layer with 16 intensity modulation units, five hidden layers featuring 80 tunable diffractive units, and an output layer with four ports. This structure allows for the encoding of multimodal data—such as visual, audio, and tactile inputs—into optical signals, which are processed through the hidden layers to yield classification results based on the output power distribution. Notably, the TDONN chip operates without the need for external compensation algorithms, achieving real-time inference directly from optical signals, thus reducing energy consumption and latency.

The training process for the TDONN chip involves preprocessing input data to extract features, followed by the adjustment of the tunable diffractive units using a customized gradient descent algorithm. The implementation of a drop-out mechanism enhances the training efficiency by deactivating non-contributory neurons, thereby accelerating convergence. Experimental results demonstrate that the TDONN chip can achieve classification accuracies of 86%, 82%, and 89% for visual, audio, and tactile modalities, respectively, with an average accuracy of 85.7%. Furthermore, the chip’s throughput is estimated at 217.6 TOPS, significantly surpassing traditional GPU capabilities, while maintaining low energy consumption. This architecture not only showcases the potential for high-performance multimodal processing but also paves the way for future developments in integrated photonic computing systems.

شارك: