120 GOPS نواة ضوئية في فيلم رقيق من نيتريد الليثيوم للتفسير والتدريب في الموقع
120 GOPS Photonic tensor core in thin-film lithium niobate for inference and in situ training

شارك:
المجلة: Nature Communications، المجلد: 15، العدد: 1
DOI: https://doi.org/10.1038/s41467-024-53261-x
PMID: https://pubmed.ncbi.nlm.nih.gov/39433733
تاريخ النشر: 2024-10-21
المؤلف: Zhongjin Lin وآخرون
الموضوع الرئيسي: الشبكات العصبية وحوسبة الخزانات

نظرة عامة

يتناول القسم التقدم في النوى الضوئية للتنسور للذكاء الاصطناعي (AI) والحوسبة العصبية، مع تسليط الضوء على إمكانياتها في الحسابات ذات الكمون المنخفض، والسرعة العالية، وكفاءة الطاقة. يقدم المؤلفون بنية جديدة متكاملة بالكامل لنواة التنسور الضوئية، والتي تتكون من جهازين تعديل رقيق من الليثيوم نيوبيوم (TFLN)، ليزر III-V، ومستقبل ضوئي بتكامل الشحن. يتيح هذا التصميم تنفيذ طبقة كاملة من الشبكة العصبية بسرعة حسابية تبلغ 120 مليار عملية في الثانية (GOPS) ويسمح بإجراء تعديلات مرنة في عدد المدخلات (fan-in) والمخرجات (fan-out).

بالإضافة إلى ذلك، تسهل نواة التنسور التدريب السريع في الموقع بتردد تحديث الوزن يبلغ 60 جيجاهرتز، مما يتيح تصنيف وتجمع صور بحجم 112 × 112 بكسل بنجاح. يتناول البحث أيضًا التحدي المتمثل في إجراء عمليات الضرب بين عددين سالبين لدعم التدريب في الموقع لمهام التجميع. يؤكد المؤلفون على أهمية الضرب السريع لمصفوفات كبيرة الحجم وتحديثات الوزن السريعة في معالجات نواة التنسور، والتي تعتبر ضرورية للتكامل الفعال للذكاء الاصطناعي عبر تطبيقات متنوعة، بما في ذلك المركبات المستقلة والبنية التحتية الذكية.

طرق

في هذه الدراسة، استخدم المؤلفون محلل شبكة متجهة (Agilent N5227A) بعرض نطاق يصل إلى 67 جيجاهرتز لتوصيف الاستجابة الكهروضوئية لمعدل تم تصنيعه عند طول موجي للاتصالات يبلغ 1310 نانومتر. تم تشغيل الجهاز بواسطة مولد موجات عشوائية (Keysight, M8194A) لأداء عملية حاصل النقطة. لأغراض التحليل المقارن، تم تنفيذ خوارزميات التعلم الآلي على وحدة المعالجة المركزية (Intel i9-9900 @ 3.10GHz)، مما يسمح بتقييم شامل لأداء المعدل. يتم تقديم مزيد من التفاصيل حول إعداد التجربة والأساليب في الملاحظة التكميلية 4.

النتائج

يقدم قسم “النتائج” من ورقة البحث النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يسلط الضوء على النتائج الهامة التي تدعم الفرضيات أو أسئلة البحث المطروحة في الدراسة. يتم عادةً توضيح البيانات من خلال أشكال متنوعة، مثل الجداول، الرسوم البيانية، أو المخططات، التي توفر تمثيلًا بصريًا للنتائج.

قد يتضمن القسم أيضًا تحليلات إحصائية تتحقق من النتائج، مثل قيم p أو فترات الثقة، مما يشير إلى موثوقية وأهمية النتائج. بشكل عام، يخدم هذا القسم لنقل الأدلة التجريبية التي تم الحصول عليها، مما يضع الأساس للنقاشات والاستنتاجات اللاحقة التي تم التوصل إليها في الورقة.

نقاش

تناقش البحث تطوير نواة التنسور الضوئية المتكاملة (IPTC) المعتمدة على تقسيم الوقت (TDM) التي تستخدم بنية مدمجة تتكون من جهازين تعديل رقيق من الليثيوم نيوبيوم (TFLN) من نوع ماخ-زندر، ليزر، ومستقبل ضوئي بتكامل الشحن. يسمح هذا التصميم المبتكر بتنفيذ طبقة كاملة من الشبكة العصبية بأحجام fan-in و fan-out قابلة للتعديل ديناميكيًا، مما يقلل بشكل كبير من عدد المكونات المطلوبة مقارنة بأنظمة تقسيم الطول الموجي التقليدية (WDM). يحقق الجهاز عمليات حاصل النقطة عالية السرعة، مما يمكّن من تحديثات الوزن السريعة والتدريب في الموقع، وهو أمر حاسم لتطبيقات التعلم الآلي الفعالة.

يظهر النموذج الأولي التكامل الناجح لهذه المكونات، محققًا سرعة حسابية تبلغ 120 مليار عملية في الثانية (GOPS) وسرعة تحديث الوزن تبلغ 60 جيجاهرتز. تشير النتائج التجريبية إلى أن IPTC يمكنه إجراء عمليات ضرب مصفوفات كبيرة الحجم بدقة عالية، محققًا انحرافًا معياريًا قدره 0.03 في قياسات حاصل النقطة، وهو ما يتجاوز الدقة المطلوبة لمهام الذكاء الاصطناعي. علاوة على ذلك، يعرض أداء الجهاز في مهام التعلم المراقب وغير المراقب، مثل تصنيف الصور والتجمع باستخدام تحليل المكونات الرئيسية (PCA)، إمكانياته للتطبيقات في العالم الحقيقي، بما في ذلك الأنظمة المستقلة. تضع قابلية توسيع البنية والتوافق مع التقنيات الضوئية الحالية الجهاز كحل واعد لشبكات الأعصاب الضوئية المستقبلية.

Journal: Nature Communications, Volume: 15, Issue: 1
DOI: https://doi.org/10.1038/s41467-024-53261-x
PMID: https://pubmed.ncbi.nlm.nih.gov/39433733
Publication Date: 2024-10-21
Author(s): Zhongjin Lin et al.
Primary Topic: Neural Networks and Reservoir Computing

Overview

The section discusses the advancements in photonic tensor cores for artificial intelligence (AI) and neuromorphic computing, highlighting their potential for low-latency, high-speed, and energy-efficient computations. The authors introduce a novel fully integrated photonic tensor core architecture, which comprises two thin-film lithium niobate (TFLN) modulators, a III-V laser, and a charge-integration photoreceiver. This design enables the implementation of an entire neural network layer with a computational speed of 120 billion operations per second (GOPS) and allows for flexible adjustments in the number of inputs (fan-in) and outputs (fan-out).

Additionally, the tensor core facilitates rapid in-situ training with a weight update frequency of 60 GHz, successfully classifying and clustering 112 × 112 pixel images. The paper also addresses the challenge of performing multiplications between two negative numbers to support in-situ training for clustering tasks. The authors emphasize the importance of high-speed, large-scale matrix-vector multiplication and rapid weight updates in tensor core processors, which are essential for the effective integration of AI across various applications, including autonomous vehicles and smart infrastructure.

Methods

In this study, the authors utilized a vector network analyzer (Agilent N5227A) with a bandwidth of up to 67 GHz to characterize the electro-optic response of a fabricated modulator at a telecom wavelength of 1310 nm. The device was driven by an arbitrary wave generator (Keysight, M8194A) to perform the dot product operation. For comparative analysis, machine learning algorithms were executed on a CPU (Intel i9-9900 @ 3.10GHz), allowing for a comprehensive evaluation of the modulator’s performance. Further details regarding the experimental setup and methodologies are provided in Supplementary Note 4.

Results

The “Results” section of the research paper presents the key findings derived from the conducted experiments or analyses. It highlights the significant outcomes that support the hypotheses or research questions posed in the study. The data is typically illustrated through various forms, such as tables, graphs, or charts, which provide a visual representation of the results.

The section may also include statistical analyses that validate the findings, such as p-values or confidence intervals, indicating the reliability and significance of the results. Overall, this section serves to convey the empirical evidence obtained, laying the groundwork for subsequent discussions and conclusions drawn in the paper.

Discussion

The research discusses the development of a Time-Division Multiplexing (TDM)-based Integrated Photonic Tensor Core (IPTC) that utilizes a compact architecture comprising two cascaded thin-film lithium niobate (TFLN) Mach-Zehnder modulators, a laser, and a charge-integration photoreceiver. This innovative design allows for the implementation of an entire neural network layer with dynamically adjustable fan-in and fan-out sizes, significantly reducing the number of required components compared to conventional Wavelength-Division Multiplexing (WDM) systems. The device achieves high-speed dot product operations, enabling rapid weight updates and in situ training, which is crucial for efficient machine learning applications.

The prototype demonstrates the successful integration of these components, achieving a computational speed of 120 Giga Operations Per Second (GOPS) and a weight update speed of 60 GHz. Experimental results indicate that the IPTC can perform large-scale matrix-vector multiplications with high precision, achieving a standard deviation of 0.03 in dot product measurements, which exceeds the precision required for artificial intelligence tasks. Furthermore, the device’s performance in supervised and unsupervised learning tasks, such as image classification and clustering using Principal Component Analysis (PCA), showcases its potential for real-world applications, including autonomous systems. The architecture’s scalability and compatibility with existing optical technologies position it as a promising solution for future photonic neural networks.

شارك: