DOI: https://doi.org/10.29026/oea.2025.250013
تاريخ النشر: 2025-01-01
المؤلف: Fujie Li وآخرون
الموضوع الرئيسي: الليزر العشوائي ووسائط التشتت
نظرة عامة
في هذا البحث، يقدم المؤلفون طريقة جديدة تستخدم شبكة تكيف المجال غير المشرف (ULDAN) لتسهيل نقل المعلومات ذات البكسل الواحد بدون تسميات عبر وسائط التشتت الديناميكية غير المعروفة. يقومون بتحسين نموذج شبكة التصوير بالبكسل الواحد (SPI) من خلال دمج وحدة محاذاة الميزات المستوحاة من التعلم العدائي، مما يسمح بنقل المعرفة بشكل فعال من نموذج مدرب مسبقًا في مجال المصدر إلى المجال المستهدف. يتيح هذا التكامل تحسين أداء SPI بدون تسميات في المجال المستهدف.
بالإضافة إلى ذلك، يجمع المؤلفون بين بنية الترميز التلقائي مع قيود سابقة فيزيائية لتصميم أقنعة تكيفية تتجاوز طرق الترميز التقليدية، مما يؤدي إلى أداء نقل صورة متفوق عند نسب ضغط أقل. تكشف محاكاةهم وتجاربهم أنه عند نسبة ضغط أقل من 0.2، تحقق الطريقة المقترحة نقل صورة عالي الدقة، مع مؤشرات تشابه هيكلي تتجاوز 0.9 عبر قنوات تشتت معقدة مختلفة. ومن الجدير بالذكر أنه في الاختبارات الواقعية التي تشمل قناة ديناميكية مضطربة تحت الماء بطول 5 أمتار، تحسن جودة التصوير بأكثر من 13 ديسيبل مقارنة بالتقنيات التقليدية. كما تظهر الطريقة أيضًا نقلًا قريبًا من عدم الفقد للفيديوهات الملونة بدقة 720×720 بمعدل 30 إطارًا في الثانية على مدى فترات طويلة، مما يبرز إمكانياتها للتطبيقات العملية في تصوير وسائط التشتت.
مقدمة
تستعرض المقدمة التقدم في تقنيات تصوير وسائط التشتت، التي تسهل تصور الأهداف المحجوبة بواسطة بيئات معقدة مثل الضباب، والاضطراب، والأنسجة البيولوجية. أظهرت الطرق التقليدية، بما في ذلك تصوير ارتباط النقاط والتشكيل الموجي، وعدًا ولكنها غالبًا ما تتطلب كواشف عالية السرعة وتكون عرضة للضوضاء. في المقابل، يستخدم تصوير البكسل الواحد (SPI) سلسلة من الأقنعة لالتقاط إشارات أحادية البعد، مما يسمح بإعادة بناء صور ثنائية الأبعاد مع حساسية وموثوقية محسنتين. ومع ذلك، لا تزال هناك تحديات في استعادة الصور بكفاءة من هذه الإشارات، خاصة بسبب التكاليف الحسابية العالية والحاجة إلى مصفوفات قياس فعالة.
لقد حسنت التطورات الأخيرة في الاستشعار المضغوط (CS) SPI من خلال الاستفادة من ندرة المشاهد، ومع ذلك لا تزال تواجه مقايضات بين دقة الإعادة والمتطلبات الحسابية. ظهرت تقنيات التعلم العميق (DL) كحل، مما يتيح إعادة بناء دقيقة عند نسب ضغط أقل؛ ومع ذلك، تتطلب عادة تدريبًا مشرفًا ومجموعات بيانات مصنفة، وهو ما قد يكون غير عملي في البيئات الديناميكية. يقدم العمل المقترح نهج SPI جديدًا بدون تسميات يستخدم شبكة تكيف المجال غير المشرف (ULDAN) لتسهيل نقل المعلومات المرنة عبر وسائط التشتت الديناميكية غير المعروفة. تتضمن هذه الطريقة تدريب نموذج SPI مسبقًا في ظروف مثالية واستخدام التعلم العدائي لمحاذاة الميزات عبر المجالات المختلفة، مما يلغي الحاجة إلى مجموعات بيانات مصنفة ويعزز من عملية نقل الصور في البيئات المعقدة. بالإضافة إلى ذلك، يتم اقتراح إطار عمل مدعوم بالفيزياء لتحسين كفاءة الترميز، مما يحقق تصويرًا عالي الجودة عند نسب ضغط منخفضة.
طرق
في هذا القسم، يصف المؤلفون تقنية جديدة للتصوير بالبكسل الواحد (SPI) بدون تسميات مصممة لتعزيز نقل الصور عبر وسائط التشتت المعقدة والمتغيرة ديناميكيًا، مثل المياه المضطربة. تكافح طرق التعلم العميق المشرف التقليدية في هذه البيئات بسبب القيود الجوهرية لتعميم الشبكات العصبية، كما هو موضح في نظرية “لا غداء مجاني”. لمعالجة ذلك، تستخدم الإطار المقترح نظرية تكيف المجال، التي تحاذي توزيعات البيانات من كل من مجالات المصدر والهدف في فضاء ميزات عالي الأبعاد مشترك. يسمح هذا المحاذاة بإعادة استخدام فعالة لشبكات إعادة بناء الصور المدربة مسبقًا، مما يسهل نقل الصور عالية الجودة حتى في الظروف الصعبة.
تظهر إعدادات التجربة فعالية طريقة شبكة التكيف المدفوعة بالتعلم غير المشرف (ULDAN) المقترحة. تشير النتائج إلى أن ULDAN تحقق درجة وفاء تتجاوز 0.9 عبر مجموعات بيانات مختلفة، متفوقة بشكل كبير على طرق النقل المباشر وقريبة من أداء طرق إعادة التدريب. كما يتحقق المؤلفون من قوة ULDAN من خلال تجارب نقل الفيديو، حيث تحقق معدل خطأ بت (BER) يبلغ 0.008 فقط ومتوسط وفاء لإطارات الفيديو المعاد بناؤها يتجاوز 0.98. تؤكد هذه النتائج على إمكانيات ULDAN لنقل البيانات بدقة عالية في التطبيقات الواقعية التي تشمل وسائط التشتت غير المعروفة، بينما تشير أيضًا إلى أن مزيدًا من تحسين آليات محاذاة الميزات قد يعزز الأداء أكثر.
نتائج
تظهر نتائج الدراسة فعالية طريقة شبكة البيانات التكيفية المعتمدة على التعلم غير المشرف (ULDAN) المقترحة لنقل الصور عبر قنوات مضطربة، خاصة بالمقارنة مع تقنيات التعلم النقل التقليدية. تكشف المحاكاة أنه بينما يحقق نموذج تصوير الفوتون الهيكلي (SPI) المدرب مسبقًا وفاءً عاليًا لإعادة البناء يبلغ 0.991 في قناة مثالية، فإن أدائه يتدهور بشكل كبير تحت الاضطراب الشديد، حيث ينخفض مؤشر التشابه الهيكلي (SSIM) إلى أقل من 0.2. بالمقابل، تحافظ طريقة ULDAN على وفاء نقل يبلغ حوالي 0.8، مما يظهر قوتها في الظروف الصعبة.
تشير التحليلات الإضافية إلى أن ULDAN تقترب من أداء طريقة إعادة التدريب النظرية، حيث تحقق حوالي 80% من فعاليتها، بينما تصل طريقة النقل المباشر إلى أقل من 30%. توضح المقارنات المرئية أنه تحت الاضطراب الشديد، تؤدي طريقة النقل المباشر إلى إعادة بناء فوضوية، بينما تنجح ULDAN في تحديد الخطوط الرئيسية لأنماط الأهداف. بالإضافة إلى ذلك، تؤكد التحليلات عالية الأبعاد باستخدام مقاييس T-SNE والمسافة فاسرشتاين أن ULDAN تحافظ على تشابه الميزات بين مجالات المصدر والهدف، وهو أمر ضروري لإعادة بناء الصور عالية الجودة. تؤكد هذه النتائج على إمكانيات ULDAN لنقل البيانات بشكل فعال في البيئات المعقدة.
نقاش
في هذا القسم، يناقش المؤلفون طريقتهم المقترحة للتعلم غير المشرف لنقل بيانات التصوير بالبكسل الواحد (SPI) بدون تسميات، والتي تستخدم نموذج إعادة بناء مدرب مسبقًا في ظروف القناة المثالية. تتضمن الطريقة عملية من خطوتين حيث يتم ترميز إشارات الكشف في فضاء ميزات عالي الأبعاد بواسطة مشفر، ثم يتم فك ترميزها مرة أخرى إلى أنماط الهدف بواسطة مفكك الترميز. لمعالجة تحدي نقص تسميات التدريب في المجال المستهدف، يقدم المؤلفون تقنية محاذاة الميزات التي تستخدم التعلم العدائي لمحاذاة الميزات من مجال المصدر مع تلك من المجال المستهدف، مما يسمح بإعادة استخدام فعالة لمفكك الترميز المدرب مسبقًا.
بالإضافة إلى ذلك، يقترح المؤلفون خوارزمية تعلم شاملة لتصميم أقنعة تكيفية تعزز أداء إعادة البناء مع الالتزام بالقيود الفيزيائية، مثل الحفاظ على توزيع ثنائي متوازن وضمان التعامد بين الأقنعة. تظهر النتائج التجريبية أن هذه الأقنعة التكيفية تتفوق بشكل كبير على الأقنعة الثنائية التقليدية عبر نسب ضغط مختلفة، محققة نقل صورة عالي الوفاء مع درجات تشابه هيكلي تتجاوز 0.9، حتى في بيئات التشتت المعقدة. تشير النتائج إلى أن الطريقة المقترحة قابلة للتوسع وقابلة للتكيف مع وسائط التشتت المتنوعة، مما يوفر تطبيقات واعدة في سيناريوهات التصوير الواقعية.
DOI: https://doi.org/10.29026/oea.2025.250013
Publication Date: 2025-01-01
Author(s): Fujie Li et al.
Primary Topic: Random lasers and scattering media
Overview
In this research, the authors introduce a novel method utilizing an unsupervised learning domain adaptation network (ULDAN) to facilitate label-free single-pixel information transmission through unknown dynamic scattering media. They enhance the single-pixel imaging (SPI) network model by incorporating a feature alignment module inspired by adversarial learning, which allows for the effective transfer of knowledge from a pre-trained model in the source domain to the target domain. This integration enables improved label-free SPI performance in the target domain.
Additionally, the authors combine an auto-encoder architecture with physical prior constraints to design adaptive masks that surpass traditional encoding methods, resulting in superior image transmission performance at lower compression ratios. Their simulations and experiments reveal that at a compression ratio below 0.2, the proposed approach achieves high-fidelity image transmission, with structural similarity indices exceeding 0.9 across various complex scattering channels. Notably, in real-world tests involving a 5 m underwater dynamic turbulent channel, the imaging quality improved by over 13 dB compared to conventional techniques. The method also successfully demonstrates near-lossless transmission of 720×720 30 fps color videos over extended durations, underscoring its potential for practical applications in scattering media imaging.
Introduction
The introduction outlines the advancements in scattering media imaging technologies, which facilitate the visualization of targets obscured by complex environments such as fog, turbulence, and biological tissues. Traditional methods, including speckle correlation imaging and wavefront shaping, have shown promise but often require high-speed detectors and are susceptible to noise. In contrast, single-pixel imaging (SPI) employs a series of masks to capture one-dimensional signals, allowing for the reconstruction of two-dimensional images with enhanced sensitivity and robustness. However, challenges remain in efficiently recovering images from these signals, particularly due to high computational costs and the need for effective measurement matrices.
Recent developments in compressed sensing (CS) have improved SPI by leveraging the sparsity of scenes, yet they still face trade-offs between reconstruction accuracy and computational demands. Deep learning (DL) techniques have emerged as a solution, enabling accurate reconstructions at lower compression ratios; however, they typically require supervised training and labeled datasets, which can be impractical in dynamic environments. The proposed work introduces a novel label-free SPI approach utilizing an unsupervised learning domain adaptation network (ULDAN) to facilitate resilient information transmission through unknown dynamic scattering media. This method involves pre-training the SPI model under ideal conditions and employing adversarial learning for feature alignment across varying domains, thereby eliminating the need for labeled datasets and enhancing the practicality of image transmission in complex environments. Additionally, a physics-informed autoencoding framework is proposed to improve coding efficiency, achieving high-quality imaging at reduced compression ratios.
Methods
In this section, the authors describe a novel unsupervised label-free single-pixel imaging (SPI) technique designed to enhance image transmission through complex and dynamically changing scattering media, such as turbulent water. Traditional supervised deep learning methods struggle in these environments due to the inherent limitations of neural network generalization, as highlighted by the “No Free Lunch” theorem. To address this, the proposed framework utilizes domain adaptation theory, which aligns data distributions from both source and target domains in a shared high-dimensional feature space. This alignment allows for the effective reuse of pre-trained image reconstruction networks, facilitating high-quality image transmission even in challenging conditions.
The experimental setup demonstrates the efficacy of the proposed unsupervised learning-driven adaptation network (ULDAN) method. Results indicate that ULDAN achieves a fidelity score exceeding 0.9 across various datasets, significantly outperforming direct transfer methods and approaching the performance of retraining methods. The authors also validate the robustness of ULDAN through video transmission experiments, achieving a bit error rate (BER) of only 0.008 and an average fidelity of reconstructed video frames exceeding 0.98. These findings underscore the potential of ULDAN for high-precision data transmission in real-world applications involving unknown scattering media, while also suggesting that further refinement of feature alignment mechanisms could enhance performance even more.
Results
The results of the study demonstrate the effectiveness of the proposed Unsupervised Learning-based Data Adaptive Network (ULDAN) method for image transmission through turbulent channels, particularly in comparison to traditional transfer learning techniques. Simulations reveal that while a pre-trained Structured Photonic Imaging (SPI) model achieves a high reconstruction fidelity of 0.991 in an ideal channel, its performance significantly deteriorates under intense turbulence, with the Structural Similarity Index (SSIM) dropping below 0.2. In contrast, the ULDAN method maintains a transmission fidelity of approximately 0.8, showcasing its robustness in challenging conditions.
Further analysis indicates that ULDAN approaches the performance of a theoretical retraining method, achieving about 80% of its efficacy, while the direct transfer method only reaches less than 30%. Visual comparisons illustrate that under intense turbulence, the direct transfer method results in chaotic reconstructions, whereas ULDAN successfully delineates the main contours of target patterns. Additionally, high-dimensional feature analysis using T-SNE and Wasserstein distance metrics confirms that ULDAN preserves feature similarity between source and target domains, which is essential for high-quality image reconstruction. These findings underscore the potential of ULDAN for effective data transmission in complex environments.
Discussion
In this section, the authors discuss their proposed unsupervised learning method for label-free single-pixel imaging (SPI) data transmission, which utilizes a pre-trained reconstruction model under ideal channel conditions. The method involves a two-step process where detection signals are encoded into a high-dimensional feature space by an Encoder, and then decoded back into target patterns by a Decoder. To address the challenge of lacking training labels in the target domain, the authors introduce a feature alignment technique that employs adversarial learning to align features from the source domain with those from the target domain, allowing for the effective reuse of the pre-trained Decoder.
Additionally, the authors propose an end-to-end learning algorithm for designing adaptive masks that enhance reconstruction performance while adhering to physical constraints, such as maintaining a balanced binary distribution and ensuring orthogonality among masks. Experimental results demonstrate that these adaptive masks significantly outperform traditional binary masks across various compression ratios, achieving high-fidelity image transmission with structural similarity scores exceeding 0.9, even in complex scattering environments. The findings indicate that the proposed method is scalable and adaptable to diverse scattering media, offering promising applications in real-world imaging scenarios.
