تقليل بصمة الذاكرة لتقنية الطباعة ثلاثية الأبعاد باستخدام Gaussian Splatting
Reducing the Memory Footprint of 3D Gaussian Splatting

المجلة: Proceedings of the ACM on Computer Graphics and Interactive Techniques، المجلد: 7، العدد: 1
DOI: https://doi.org/10.1145/3651282
تاريخ النشر: 2024-05-11
المؤلف: Panagiotis Papantonakis وآخرون
الموضوع الرئيسي: الرؤية المتقدمة والتصوير

نظرة عامة

في هذه الورقة، نقدم تقنية شاملة وفعالة لتقليل الذاكرة لتدفق الرسوميات ثلاثية الأبعاد (3DGS). تتضمن طريقتنا طريقة تقليل بدائية مدركة للدقة تقلل بشكل فعال عدد البدائيات إلى النصف، وآلية تعديل تكيفية لتحديد العدد الأمثل من نطاقات التوافقيات الكروية (SH) لكل بدائية، واستراتيجية تشفير قائمة على كتاب الرموز. هذه الابتكارات مجتمعة تؤدي إلى تقليل ملحوظ في استخدام الذاكرة بمقدار 27 مرة مع تعزيز سرعة العرض بمقدار 1.7 مرة.

نحن نتحقق من طريقتنا ضمن إطار عمل تدفق يستخدم تنفيذ WebGL، محققين تقليلاً في وقت التنزيل بمقدار 20-30 مرة وزيادة تقريبية في سرعة العرض بمقدار ثلاثة أضعاف. تمثل هذه العمل أول حل لتدفق/المحمول 3DGS يحافظ على جودة بصرية عالية، مما يعالج قيداً حاسماً في هذا المجال. تبرز طريقتنا كأكثر الحلول التنافسية للتخزين غير المتطاير (NVS) عبر أبعاد السرعة والجودة وكفاءة الذاكرة. قد تستكشف الأبحاث المستقبلية مزيداً من التخفيضات في عدد البدائيات واستراتيجيات لمنع الكثافة الزائدة الأولية، مستفيدة من البيانات المدفوعة، مثل الإشراف على العمق، كما تشير الاختبارات الأولية.

مقدمة

تناقش مقدمة هذه الورقة البحثية التقدم في توليد المناظر الجديدة (NVS)، والذي يتضمن إنشاء مناظر ثلاثية الأبعاد جديدة من الصور أو مقاطع الفيديو. على مدار الـ 25 عاماً الماضية، تطور هذا المجال من العرض القائم على الصور التقليدية إلى العرض العصبي، مع مساهمات كبيرة من مجالات الإشعاع العصبي (NeRFs) التي تستخدم شبكة متعددة الطبقات (MLP) لتمثيل المشهد بشكل ضمني. بينما حققت NeRFs جودة بصرية عالية، إلا أنها عانت من أوقات عرض بطيئة ومتطلبات ذاكرة عالية. استهدفت الطرق الحديثة، بما في ذلك 3D Gaussian Splatting (3DGS)، تحقيق توازن بين السرعة والجودة وكفاءة الذاكرة، لكن تنفيذ 3DGS الأصلي لا يزال يقدم بصمة ذاكرة كبيرة (700MB-1.2GB).

لمعالجة هذه التحديات، يقترح المؤلفون نهجاً فعالاً من حيث الذاكرة يقلل من حجم التمثيل بمقدار 27× بينما يزيد من سرعة العرض بمقدار 1.7×. يتم تحقيق ذلك من خلال ثلاث مساهمات رئيسية: (1) طريقة تقليم بدائية مدركة للدقة تقضي على حوالي 60% من البدائيات ثلاثية الأبعاد غير الضرورية، (2) تقنية تعديل تكيفية لنطاقات التوافقيات الكروية (SH) التي تحسن عدد النطاقات المستخدمة لكل بدائية، و(3) طريقة تشفير قائمة على كتاب الرموز مدمجة مع تمثيل نصف عائم للتخزين الفعال. تتضمن الورقة تقييمًا شاملاً لهذه الطرق عبر مجموعات بيانات قياسية، مما يظهر تحسينات كبيرة في تقليل الذاكرة وسرعة العرض، خاصة للتطبيقات المحمولة.

النتائج

في هذا القسم، يقدم المؤلفون تقييمًا شاملاً لطرقهم المطبقة على مجموعات بيانات قياسية، بما في ذلك MipNeRF360 وDeep Blending وTanks&Temples، باستخدام وحدة معالجة الرسوميات NVIDIA RTX A6000. تظهر النتائج أن تقنية تقليل البدائيات الخاصة بهم تقلل بشكل كبير من بصمة الذاكرة إلى ما بين 32% و52% من الطريقة الأصلية مع الحفاظ على تأثير ضئيل على نسبة الإشارة إلى الضوضاء القصوى (PSNR)، والتي تتراوح من -0.32 إلى +0.16 ديسيبل. يتم تعويض هذا التدهور الطفيف في الجودة البصرية من خلال تحسينات في التعميم بسبب تأثيرات تنظيم تقليم البدائيات والتوافقيات الكروية (SH).

تحقق مساهمات تقليم البدائيات وSH، جنبًا إلى جنب مع التشفير، تقليلاً متوسطًا بمقدار 27× في حجم تمثيل 3DGS عبر جميع مجموعات البيانات. يبرز المؤلفون أن دمج كتب الرموز والتعامل مع نوع نصف العائم يمكن أن يحسن استهلاك VRAM بشكل أكبر. تؤدي طريقتهم إلى تحسين كبير في كفاءة التدفق، حيث تقلل أوقات التحميل من 120 ثانية إلى 5 ثوانٍ فقط بينما تزيد من معدلات الإطارات من 16 إطارًا في الثانية إلى 45 إطارًا في الثانية، مما يحقق تسريعًا بمقدار 24× في وقت التنزيل. تشير المقارنات البصرية إلى أن تدهور الجودة يكاد يكون غير ملحوظ، خاصة على الشاشات الأصغر، مما يضع نهجهم في موقع متميز مقارنةً بأساليب الوقت الحقيقي الأخرى المتطورة.

المناقشة

في مناقشة توليد المناظر الجديدة (NVS)، تستعرض الورقة تطور الخوارزميات من العرض القائم على الصور (IBR) إلى التقدمات الحديثة في مجالات الإشعاع العصبي (NeRFs). تستخدم طرق IBR وكلاء هندسيين لإعادة إسقاط بيانات العرض المدخلة ولكن تواجه تحديات مع العيوب وقابلية توسيع الذاكرة مع زيادة تعقيد المشهد. قدمت NeRFs تمثيلاً حجميًا قابلاً للاشتقاق، مما يسمح بالتحسين عبر الانحدار العشوائي (SGD). بينما توفر NeRFs تمثيلاً مضغوطًا، إلا أنها غالبًا ما تتطلب تدريبًا مكثفًا وأوقات عرض طويلة. حاولت الطرق الحديثة تعزيز السرعة وكفاءة الذاكرة، بما في ذلك الشبكات الفراغية والهياكل النادرة، ومع ذلك لا تزال هناك تنازلات بين الجودة والسرعة واستهلاك الذاكرة.

تسلط الورقة الضوء على طريقة 3D Gaussian Splatting (3DGS)، التي تقدم تمثيلاً نادرًا لمجالات الإشعاع، محققة سرعات عرض عالية مع الحفاظ على الجودة. ومع ذلك، تشير أيضًا إلى أن بصمة الذاكرة لـ 3DGS أعلى بكثير من الطرق السابقة. يقترح المؤلفون استراتيجيات لتقليل هذا الاستخدام للذاكرة، بما في ذلك تحديد وتقليم البدائيات الزائدة، وتكييف عدد نطاقات التوافقيات الكروية (SH) بناءً على الحاجة، واستخدام طريقة ضغط قائمة على كتاب الرموز. تهدف هذه الاستراتيجيات مجتمعة إلى الحفاظ على مزايا السرعة والجودة لـ 3DGS مع تحقيق تقليل كبير في بصمة الذاكرة، مع الإشارة إلى أن النتائج التجريبية تشير إلى تقليل إجمالي للذاكرة بحوالي 87% مع تدهور طفيف في الجودة.

Journal: Proceedings of the ACM on Computer Graphics and Interactive Techniques, Volume: 7, Issue: 1
DOI: https://doi.org/10.1145/3651282
Publication Date: 2024-05-11
Author(s): Panagiotis Papantonakis et al.
Primary Topic: Advanced Vision and Imaging

Overview

In this paper, we introduce a comprehensive and efficient memory reduction technique for 3D Graphics Streaming (3DGS). Our approach incorporates a resolution-aware primitive reduction method that effectively halves the number of primitives, an adaptive adjustment mechanism to determine the optimal number of spherical harmonics (SH) bands for each primitive, and a codebook-based quantization strategy. These innovations collectively yield a remarkable 27-fold reduction in memory usage while enhancing rendering speed by 1.7 times.

We validate our method within a streaming framework utilizing a WebGL implementation, achieving a 20-30 times reduction in download time and approximately a threefold increase in rendering speed. This work represents the first streaming/mobile 3DGS solution that maintains high visual quality, addressing a critical limitation in the field. Our method stands out as the most competitive non-volatile storage (NVS) solution across the dimensions of speed, quality, and memory efficiency. Future research may explore further reductions in primitive counts and strategies to prevent initial over-densification, potentially leveraging data-driven priors, such as depth supervision, as indicated by preliminary tests.

Introduction

The introduction of this research paper discusses advancements in novel view synthesis (NVS), which involves generating new 3D views from photographs or videos. Over the past 25 years, the field has evolved from traditional image-based rendering to neural rendering, with significant contributions from Neural Radiance Fields (NeRFs) that utilize a Multi-Layer Perceptron (MLP) for implicit scene representation. While NeRFs achieved high visual quality, they suffered from slow rendering times and high memory requirements. Recent methods, including 3D Gaussian Splatting (3DGS), have aimed to balance speed, quality, and memory efficiency, but the original 3DGS implementation still presented a substantial memory footprint (700MB-1.2GB).

To address these challenges, the authors propose a memory-efficient approach that reduces the representation size by 27× while increasing rendering speed by 1.7×. This is achieved through three main contributions: (1) a resolution-aware primitive pruning method that eliminates approximately 60% of unnecessary 3D primitives, (2) an adaptive adjustment technique for Spherical Harmonic (SH) bands that optimizes the number of bands used per primitive, and (3) a codebook quantization method combined with half-float representation for efficient storage. The paper includes a comprehensive evaluation of these methods across standard datasets, demonstrating significant improvements in memory reduction and rendering speed, particularly for mobile applications.

Results

In this section, the authors present a comprehensive evaluation of their methods applied to standard datasets, including MipNeRF360, Deep Blending, and Tanks&Temples, utilizing an NVIDIA RTX A6000 GPU. The results demonstrate that their primitive reduction technique significantly decreases the memory footprint to between 32% and 52% of the original method while maintaining a minimal impact on Peak Signal-to-Noise Ratio (PSNR), ranging from -0.32 to +0.16 dB. This slight degradation in visual quality is offset by improvements in generalization due to the regularization effects of primitive and spherical harmonics (SH) culling.

The contributions of primitive and SH culling, along with quantization, collectively achieve an average reduction of 27× in the size of the 3DGS representation across all datasets. The authors highlight that integrating codebooks and half-float type handling could further optimize VRAM consumption. Their method results in a substantial improvement in streaming efficiency, reducing loading times from 120 seconds to just 5 seconds while increasing frame rates from 16 FPS to 45 FPS, thus achieving a 24× speedup in download time. Visual comparisons indicate that the quality degradation is virtually imperceptible, particularly on smaller displays, positioning their approach favorably against other state-of-the-art real-time methods.

Discussion

In the discussion of Novel View Synthesis (NVS), the paper reviews the evolution of algorithms from Image-Based Rendering (IBR) to recent advancements in Neural Radiance Fields (NeRFs). IBR methods utilize geometry proxies for reprojecting input view data but face challenges with artifacts and memory scalability as scene complexity increases. NeRFs introduced a differentiable volumetric representation, allowing for optimization via Stochastic Gradient Descent (SGD). While NeRFs provide a compact representation, they often require extensive training and rendering times. Recent approaches have attempted to enhance speed and memory efficiency, including voxel grids and sparse structures, yet trade-offs between quality, speed, and memory consumption persist.

The paper highlights the 3D Gaussian Splatting (3DGS) method, which offers a sparse representation of radiance fields, achieving high rendering speeds while maintaining quality. However, it also notes that 3DGS’s memory footprint is significantly higher than previous methods. The authors propose strategies to reduce this memory usage, including identifying and pruning redundant primitives, adapting the number of spherical harmonics (SH) bands based on necessity, and employing a codebook compression method. These strategies collectively aim to maintain the speed and quality advantages of 3DGS while achieving a substantial reduction in memory footprint, with experimental results indicating an overall memory reduction of approximately 87% with minimal quality degradation.