DOI: https://doi.org/10.1186/s43593-026-00127-y
تاريخ النشر: 2026-03-18
المؤلف: Mingcheng Luo وآخرون
الموضوع الرئيسي: الشبكات العصبية وحوسبة الخزانات
نظرة عامة
تناقش هذه الفقرة التقدم في الشبكات العصبية الضوئية (ONNs)، مع تسليط الضوء على إمكانياتها في الذكاء الاصطناعي منخفض الكمون وذو الكفاءة في استهلاك الطاقة. ومع ذلك، تواجه توسيع هذه الشبكات تحديات بسبب المتطلبات الحسابية لتدريب ONNs الكبيرة وحساسية المكونات الضوئية لعيوب التصنيع. يقدم المؤلفون آلة تعلم ضوئية تعتمد على السطح الميتا، تعالج هذه القضايا من خلال استخدام شبكة عصبية واسعة للغاية (NN) تحتوي على 41 مليون معلمة ضوئية. ومن المRemarkably، يظهرون أن سطح ميتا ثابت وغير مدرب يمكن أن يقترب بفعالية من نموذج مدرب بالكامل، مع إدارة الفروقات المتبقية بواسطة واجهة رقمية مدمجة تحتوي فقط على $10^2$ إلى $10^4$ معلمات قابلة للتدريب.
تحقق النظام الهجين المطور دقة تنافسية عبر ست مهام للرؤية الآلية، مقارنة بالنماذج الرائدة مثل ResNet و Vision Transformer. بالإضافة إلى ذلك، يظهر مقاومة لأخطاء التصنيع والانحراف، ويعمل بفعالية تحت كل من الإضاءة المتماسكة وغير المتماسكة. لا تعزز هذه الطريقة المبتكرة قابلية التوسع فحسب، بل توفر أيضًا مسارًا قابلاً للتطبيق لتنفيذ حوسبة الذكاء الاصطناعي الضوئي عالية الأداء على نطاق واسع.
مقدمة
تناقش مقدمة هذه الورقة البحثية التأثير التحويلي للشبكات العصبية (NNs) عبر مجالات مختلفة، مثل اكتشاف الأدوية والتشخيص الطبي. ومع ذلك، أدت زيادة تعقيد المشكلات إلى ارتفاع سريع في التكاليف المرتبطة بتدريب واستنتاج نماذج NN، والتي تتضاعف كل شهرين، متجاوزة التقدم في تكنولوجيا CMOS. تقدم الشبكات العصبية الضوئية (ONNs) بديلاً واعدًا بسبب توازيها الفطري، الذي يمكن أن يعزز كفاءة الطاقة والسرعة. على الرغم من إمكانياتها، تكافح تطبيقات ONN الحالية مع توسيع عدد المعلمات، خاصة في التكوينات ثنائية الأبعاد (2D)، التي تحدها حجم وتعقيد المكونات الضوئية.
للتغلب على هذه القيود، تقترح الورقة بنية هجينة تستخدم سطح ميتا ضوئي كبير الحجم يتكون من 41 مليون ميتا-ذرة، مستوحاة من آلات التعلم المتطرفة (ELMs). تتميز هذه البنية بطبقة مخفية واسعة لاستخراج الميزات عالية السعة وNN رقمي مدمج للتعلم الخاص بالمهام. من خلال تثبيت الواجهة الضوئية وتدريب الطبقة الرقمية فقط، يحقق النظام دقة تنافسية عبر مهام متنوعة، بما في ذلك التطبيقات الواقعية مثل تشخيص سرطان الثدي. يظهر النظام المقترح مزايا كبيرة في قابلية التوسع، حيث يعمل بكفاءة تحت ظروف مختلفة ويتحمل عيوب التصنيع، مما يمهد الطريق لحوسبة الذكاء الاصطناعي الضوئي عالية الأداء.
طرق
في هذا القسم، يحدد المؤلفون الطرق التجريبية المستخدمة لإظهار الأداء العالي والمرونة وقابلية التوسع لنظام الرؤية الآلية الخاص بهم، والذي يُشار إليه باسم MOLM. شملت المقارنات مع ثلاثة نماذج بارزة في التعلم العميق: ResNet-50، وهو شبكة عصبية تلافيفية تقليدية (CNN) تحتوي على حوالي 23.5 مليون معلمة؛ نموذج Segment Anything (SAM)، وهو نموذج تقسيم متقدم يحتوي على 93.7 مليون معلمة؛ و Vision Transformer (ViT)، الذي يحتوي على أكثر من 85.8 مليون معلمة. يستخدم الإعداد التجريبي، الموضح في الشكل 2a، جهاز تعديل الضوء المكاني (SLM) لتوليد الصور المدخلة، والتي تتم معالجتها بواسطة سطح ميتا ثم يتم التقاطها بواسطة كاميرا رقمية CMOS.
تؤكد المنهجية على استخدام شريحة سطح ميتا واحدة عبر مهام مختلفة، مع تدريب شبكة عصبية رقمية (NN) فقط في الخلفية لتطبيقات مختلفة. لا تعرض هذه الطريقة فقط تكيف النظام مع المهام المتنوعة دون الحاجة إلى إعادة تدريب السطح الميتا، بل تسلط الضوء أيضًا على المزايا العملية في التنفيذ. تشمل المهام المختارة للعرض: (1) معالجة شبكة عصبية رقمية فائقة الصغر لتصوير COVID-19، (2) التعامل مع الصور عالية الدقة لاكتشاف الأمراض الصدرية في لقطة واحدة، و(3) تصنيف واكتشاف متعدد المهام تم تحقيقه من خلال إعادة تصميم شبكة رقمية صغيرة.
نتائج
يقدم قسم “النتائج” في الورقة البحثية النتائج الرئيسية المستمدة من التجارب أو التحليلات التي تم إجراؤها. يسلط الضوء على النتائج المهمة التي ظهرت من البيانات، بما في ذلك التحليلات الإحصائية، والاتجاهات الملحوظة، وأي علاقات تم تحديدها. عادةً ما تدعم النتائج الأشكال والجداول أو المعادلات ذات الصلة التي توضح النتائج بوضوح.
قد يناقش القسم أيضًا تداعيات هذه النتائج فيما يتعلق بفرضية البحث أو الأهداف، مع التأكيد على كيفية مساهمتها في المعرفة الحالية. يتم الإشارة إلى أي نتائج غير متوقعة أو شذوذ، مع تقديم تفسيرات أو فرضيات محتملة لهذه الملاحظات. بشكل عام، يخدم هذا القسم لتوفير نظرة شاملة على الأدلة التجريبية التي تم جمعها خلال الدراسة.
مناقشة
تدمج آلة التعلم الضوئية المعتمدة على السطح الميتا (MOLM) واجهة ضوئية قائمة على السطح الميتا مع شبكة عصبية رقمية مدمجة (NN) لتعزيز مهام تصنيف الصور. تتكون الواجهة الضوئية من 41 مليون ميتا-ذرة سيليكونية أسطوانية، والتي تعدل الطور الضوئي بناءً على توزيع غاوسي، مما يمكّن من استخراج الميزات بفعالية من خلال تحويل فورييه الذي تسهلّه عدسة ضوئية. يظهر النظام تحسينات كبيرة في دقة التصنيف، حيث يحقق 99.2% على مجموعة بيانات MNIST و91.6% على مجموعة بيانات CIFAR-10، من خلال الاستفادة من ارتفاع إنتروبيا المعلومات والمكونات عالية التردد الغنية التي تولدها السطح الميتا. من الجدير بالذكر أن أداء MOLM يقترب من أداء الشبكات المدربة بالكامل مع جزء من المعلمات، مما يبرز فعالية المعلمات الضوئية في تعزيز دقة النموذج.
في التطبيقات العملية، يتفوق MOLM في مهام التصوير الطبي، حيث يحقق دقة 98.0% في تصنيف تصوير COVID-19 باستخدام 192 وزنًا رقميًا فقط، و85.4% دقة في اكتشاف الأمراض الصدرية من صور الأشعة السينية عالية الدقة. يتم إظهار تكيف النظام بشكل أكبر في اكتشاف النزيف داخل الجمجمة، حيث يحقق دقة وحساسية عالية. بالإضافة إلى ذلك، يسمح تصميم MOLM بمعالجة قابلة للتوسع وفعالة للصور الكاملة متعددة الجيجابكسل لتشخيص السرطان، مما يقلل بشكل كبير من وقت الاستدلال مقارنة بالطرق التقليدية. تؤكد النتائج على إمكانيات MOLM في التطبيقات الزمنية الحقيقية، مع التأكيد على متانته ضد تباينات التصنيع وتوافقه مع مصادر الضوء غير المتماسكة، مما يوسع من قابليته للتطبيق عبر مجالات مختلفة.
DOI: https://doi.org/10.1186/s43593-026-00127-y
Publication Date: 2026-03-18
Author(s): Mingcheng Luo et al.
Primary Topic: Neural Networks and Reservoir Computing
Overview
The section discusses advancements in optical neural networks (ONNs), highlighting their potential for low-latency and energy-efficient artificial intelligence. However, scaling these networks faces challenges due to the computational demands of training large ONNs and the sensitivity of optical components to fabrication imperfections. The authors present a metasurface-based optical learning machine that addresses these issues by utilizing an ultra-wide neural network (NN) with 41 million optical parameters. Remarkably, they demonstrate that a fixed, untrained metasurface can effectively approximate a fully trained model, with residual discrepancies managed by a compact digital backend containing only $10^2$ to $10^4$ trainable parameters.
The hybrid system developed achieves competitive accuracy across six machine vision tasks, comparable to state-of-the-art models such as ResNet and Vision Transformer. Additionally, it exhibits resilience to fabrication errors and misalignment, operating effectively under both coherent and incoherent illumination. This innovative approach not only enhances scalability but also provides a viable pathway for the implementation of large-scale, high-performance optical AI computing.
Introduction
The introduction of this research paper discusses the transformative impact of neural networks (NNs) across various fields, such as drug discovery and medical diagnostics. However, the increasing complexity of problems has led to a rapid rise in the costs associated with training and inferring NN models, which is doubling every two months, outpacing advancements in CMOS technology. Optical neural networks (ONNs) present a promising alternative due to their inherent parallelism, which can enhance energy efficiency and speed. Despite their potential, current ONN implementations struggle with scaling parameter numbers, particularly in two-dimensional (2D) configurations, which are limited by the size and complexity of optical components.
To overcome these limitations, the paper proposes a hybrid architecture utilizing a large-scale optical metasurface composed of 41 million meta-atoms, inspired by extreme learning machines (ELMs). This architecture features a wide hidden layer for high-capacity feature extraction and a compact digital NN for task-specific learning. By fixing the optical frontend and training only the digital layer, the system achieves competitive accuracy across diverse tasks, including real-world applications like breast cancer diagnosis. The proposed system demonstrates significant scalability advantages, operating efficiently under various conditions and tolerating fabrication imperfections, thereby paving the way for high-performance optical AI computing.
Methods
In this section, the authors outline the experimental methods employed to demonstrate the high performance, versatility, and scalability of their machine vision system, referred to as MOLM. The benchmarking involved comparisons with three prominent deep learning models: ResNet-50, a classical convolutional neural network (CNN) with approximately 23.5 million parameters; the Segment Anything Model (SAM), a state-of-the-art segmentation model with 93.7 million parameters; and the Vision Transformer (ViT), which has over 85.8 million parameters. The experimental setup, depicted in Figure 2a, utilizes a spatial light modulator (SLM) to generate input images, which are processed by a metasurface and subsequently captured by a CMOS digital camera.
The methodology emphasizes the use of a single metasurface chip across various tasks, with only the digital neural network (NN) at the backend being trained for different applications. This approach not only showcases the system’s adaptability to diverse tasks without the need for retraining the metasurface but also highlights practical advantages in implementation. The tasks selected for demonstration include: (1) ultra-compact digital neural network processing for COVID-19 radiography, (2) high-resolution image handling for thoracic disease detection in a single shot, and (3) multitask classification and detection achieved by redesigning a small digital NN.
Results
The “Results” section of the research paper presents key findings derived from the conducted experiments or analyses. It highlights the significant outcomes that emerged from the data, including statistical analyses, observed trends, and any correlations identified. The results are typically supported by relevant figures, tables, or equations that illustrate the findings clearly.
The section may also discuss the implications of these results in relation to the research hypothesis or objectives, emphasizing how they contribute to the existing body of knowledge. Any unexpected findings or anomalies are noted, along with potential explanations or hypotheses for these observations. Overall, this section serves to provide a comprehensive overview of the empirical evidence gathered during the study.
Discussion
The proposed metasurface-based optical learning machine (MOLM) integrates a metasurface optical frontend with a compact digital neural network (NN) to enhance image classification tasks. The optical frontend comprises 41 million silicon cylindrical meta-atoms, which modulate optical phase based on a Gaussian distribution, enabling effective feature extraction through a Fourier transform facilitated by an optical lens. The system demonstrates significant improvements in classification accuracy, achieving 99.2% on the MNIST dataset and 91.6% on the CIFAR-10 dataset, by leveraging the high information entropy and rich high-frequency components generated by the metasurface. Notably, the MOLM’s performance approaches that of fully trained networks with a fraction of the parameters, highlighting the efficacy of the optical parameters in enhancing model accuracy.
In practical applications, the MOLM excels in medical imaging tasks, achieving 98.0% accuracy in COVID-19 radiography classification using only 192 digital weights, and 85.4% accuracy in detecting thoracic diseases from high-resolution X-ray images. The system’s adaptability is further demonstrated in intracranial hemorrhage detection, where it achieves high precision and sensitivity. Additionally, the MOLM’s architecture allows for scalable and efficient processing of multi-gigapixel whole slide images for cancer diagnosis, significantly reducing inference time compared to traditional methods. The findings underscore the potential of the MOLM in real-time applications, emphasizing its robustness against fabrication variances and compatibility with incoherent light sources, thus broadening its applicability across various domains.
