DOI: https://doi.org/10.1111/mice.13201
تاريخ النشر: 2024-04-14
المؤلف: Yixiong Jing وآخرون
الموضوع الرئيسي: المسح ثلاثي الأبعاد والتراث الثقافي
نظرة عامة
تقدم ورقة البحث نهجًا جديدًا لتجزئة سحابة النقاط ثلاثية الأبعاد من خلال الاستفادة من بنية المحولات، التي أظهرت نجاحًا كبيرًا في معالجة اللغة الطبيعية (NLP). يقدم المؤلفون المحول الجغرافي المحلي، وهو وحدة تعمل على تحسين آلية الانتباه من خلال التركيز على المناطق المحلية، مما يقلل من استهلاك الذاكرة المرتبط عادةً بالمحولات التقليدية. يعالج هذا التصميم التحدي المتمثل في فقدان المعلومات الذي ينشأ من أخذ عينات فرعية من سحب النقاط أثناء التجزئة.
بالإضافة إلى ذلك، تقترح الورقة المحول الجغرافي العالمي لالتقاط العلاقات العالمية بين الميزات عند أدنى دقة. يتم التحقق من فعالية هذه البنية الجديدة باستخدام مجموعة بيانات جسر من الحجر، التي تم تطويرها بالتعاون مع شبكة التجزئة السابقة لهم، BridgeNet. تشير النتائج إلى أن الشبكة المحدثة، BridgeNetv2، تتفوق على سابقتها في جميع مقاييس التقييم مع الحفاظ على ملف خفيف وفعال من حيث الذاكرة، مما يجعلها مناسبة لمهام تجزئة سحابة النقاط على نطاق واسع في الهندسة المدنية.
مقدمة
تناقش مقدمة الورقة الزيادة في اعتماد تقنية المسح بالليزر لجمع البيانات ثلاثية الأبعاد في الهندسة المدنية، مع التأكيد على قدرتها على توليد سحب نقاط مفصلة تسهل التحليل الهيكلي والتفتيش. تعتبر تجزئة هذه السحب إلى مكونات هيكلية متميزة أمرًا حيويًا، خاصة لتقييم جسور الأقواس الحجرية، التي غالبًا ما تفتقر إلى توثيق شامل للبناء. تعتبر طرق التجزئة اليدوية التقليدية مستهلكة للوقت بسبب الحجم الكبير لسحب النقاط، مما يدفع لاستكشاف تقنيات التعلم العميق (DL) لأتمتة هذه العملية. أظهرت أساليب DL، مثل PointNet وخلفائها، وعدًا في تجزئة سحابة النقاط من خلال القضاء على الحاجة إلى ميزات محددة مسبقًا، على الرغم من أن التحديات لا تزال قائمة فيما يتعلق بالكفاءة الحسابية والأداء على مجموعات البيانات غير المتوازنة.
تقدم الورقة BridgeNetv2، وهي بنية جديدة تدمج كتلة قائمة على المحولات لاستخراج الميزات المحلية من سحب النقاط. من خلال استخدام نهج الجيران الأقرب (k-NN) لآليات الانتباه، تهدف BridgeNetv2 إلى تقليل استهلاك الذاكرة مع الحفاظ على الدقة. يتم تدريب البنية على مجموعة بيانات اصطناعية من جسور الأقواس الحجرية ويتم اختبارها على سحب نقاط من العالم الحقيقي، مما يعالج ندرة البيانات المناسبة لتدريب نماذج التجزئة في الهندسة المدنية. تم تصميم الطريقة المقترحة لتكون قابلة للتكيف مع أنواع مختلفة من الهياكل الهندسية المدنية على نطاق واسع، مع توفير تفاصيل البنية وتقييمات الأداء في الأقسام التالية.
الطرق
تصف قسم المنهجية بنية BridgeNetv2، التي تستخدم إطار عمل تقليدي للتشفير وفك التشفير لمعالجة سحابة النقاط. يتكون مشفر الميزات من أربع مراحل تستخرج تدريجيًا وتقلل من الميزات باستخدام أخذ عينات عشوائية (RS) بمعدل أخذ عينات قدره 4، مشابهًا للنهج في RandLA-Net (Hu et al.، 2020). يتم تحويل سحابة النقاط المدخلة، التي كانت في البداية بحجم $(N، 3)$، إلى تمثيل عالي الأبعاد قدره $(N، 8)$ من خلال تضمين المدخلات. تتضمن كل مرحلة طبقتين لاستخراج الميزات (FE)، مشابهتين لتلك الموجودة في BridgeNet الأصلية (Jing et al.، 2022)، التي استخدمت تقليديًا تجميع الميزات المجاورة باستخدام التجميع الأقصى. ومع ذلك، فإن هذه الطريقة تخاطر بفقدان المعلومات الهندسية ذات التردد المنخفض.
لمعالجة هذه القيود، تقدم BridgeNetv2 محول جغرافي محلي بين كل كتلة FE، والذي يحسب الانتباه المحلي على سحب النقاط قبل وبعد RS. تمكن هذه الوحدة من تجميع الميزات من النقاط التي يتم التخلص منها أثناء RS، مما يحافظ على المعلومات الهندسية من الطبقات السابقة. بالإضافة إلى ذلك، يتم تنفيذ محول جغرافي عالمي لحساب الانتباه الذاتي على الميزات ذات الدقة الأدنى، مما يلتقط الاعتمادات بعيدة المدى. التحسينات في BridgeNetv2 هي بشكل أساسي إدراج المحولات الجغرافية المحلية والعالمية، بينما تظل البنية العامة متسقة مع تلك الخاصة بـ BridgeNet. للحصول على تفاصيل شاملة حول البنية الأصلية، يتم توجيه القراء إلى Jing et al. (2022).
النتائج
في قسم النتائج، يتم مقارنة أداء طرق التجزئة الدلالية المختلفة، بما في ذلك PointNet++ وRandLA-Net وBridgeNet وBridgeNetv2 المقترحة. تظهر BridgeNetv2 أداءً متفوقًا عبر جميع المقاييس التي تم تقييمها، حيث تحقق زيادة بنسبة 5.5% في الدقة المتوسطة ($\text{mean\_Pr}$) وزيادة بنسبة 9.1% في متوسط التقاطع على الاتحاد ($\text{mIoU}$) مقارنة بسابقتها، BridgeNet. ومن الجدير بالذكر أن التحسينات في دقة التجزئة كانت ملحوظة بشكل خاص في الفئات الأخرى غير جدار السند، حيث حسنت BridgeNetv2 بشكل كبير من تجزئة فئة الفتحة، التي تم تمثيلها بشكل ناقص في مجموعة البيانات.
تظهر التقييمات البصرية لتوقعات BridgeNetv2 توافقًا قويًا مع الحقيقة الأرضية، يتميز بحدود محددة جيدًا بين المكونات. ومع ذلك، تم ملاحظة تصنيفات خاطئة في مناطق معينة، تعزى إلى التباينات بين بيانات التدريب الاصطناعية والهندسيات الواقعية، بالإضافة إلى التحديات التي تطرحها عدم توازن البيانات. يقترح المؤلفون أن تحسين عملية المحاكاة لمعالجة هذه الاختلالات بشكل أفضل يمكن أن يؤدي إلى تحسينات إضافية في دقة التجزئة.
المناقشة
يقدم قسم المناقشة في الورقة بنية ووظيفة المحولات الجغرافية المحلية والعالمية المدمجة في BridgeNetv2، وهو نموذج تعلم عميق جديد لتجزئة سحابة النقاط لجسور الأقواس الحجرية. يستخدم المحول الجغرافي المحلي آلية انتباه محلية، تحسب الانتباه بناءً على ميزات أقرب الجيران في سحب النقاط، مما يحافظ على المعلومات الهندسية أثناء أخذ العينات الفرعية. يتم تقديم تضمين موضعي غير قابل للدوران، يسمى تضمين الجغرافيا المحلي، لتعزيز حساسية النموذج للميزات الهندسية مع ضمان المتانة ضد التحولات الصلبة. في المقابل، يحسب المحول الجغرافي العالمي الانتباه الذاتي عبر الميزات عند دقة أقل، مما يلتقط العلاقات السياقية العالمية.
توضح الورقة أيضًا تفاصيل التنفيذ، بما في ذلك استخدام مُحسِّن آدم، وتقنيات زيادة البيانات، ومقاييس الأداء مثل الدقة العامة ومتوسط التقاطع على الاتحاد (mIoU). تكشف اختبارات الإزالة أن قيمة $k$ المتسقة عبر الطبقات تحسن الأداء، بينما يتفوق تضمين الموضع غير القابل للدوران المقترح بشكل كبير على الطرق التقليدية. على الرغم من الزيادة المتواضعة في حجم المعلمات واستهلاك الذاكرة، تظهر BridgeNetv2 أداءً متفوقًا مقارنة بالنماذج الحالية مثل PointNet++ وRandLA-Net، مما يشير إلى إمكاناتها لتطبيقات أوسع في مهام الهندسة المدنية. تهدف الأعمال المستقبلية إلى توسيع قابلية تطبيق النموذج على أنواع هياكل مختلفة، مما يتطلب تطوير مجموعات بيانات واسعة النطاق.
DOI: https://doi.org/10.1111/mice.13201
Publication Date: 2024-04-14
Author(s): Yixiong Jing et al.
Primary Topic: 3D Surveying and Cultural Heritage
Overview
The research paper presents a novel approach to 3D point cloud segmentation by leveraging the Transformer architecture, which has demonstrated significant success in natural language processing (NLP). The authors introduce the Local Geo-Transformer, a module that optimizes the attention mechanism by focusing on local regions, thereby reducing the memory consumption typically associated with traditional Transformers. This design addresses the challenge of information loss that arises from subsampling point clouds during segmentation.
Additionally, the paper proposes the Global Geo-Transformer to capture global relationships among features at the lowest resolution. The effectiveness of this new architecture is validated using a masonry bridge dataset, which was developed in conjunction with their previous segmentation network, BridgeNet. The results indicate that the updated network, BridgeNetv2, surpasses its predecessor in all evaluation metrics while maintaining a lightweight and memory-efficient profile, making it suitable for large-scale point cloud segmentation tasks in civil engineering.
Introduction
The introduction of the paper discusses the increasing adoption of laser scanning technology for 3D data collection in civil engineering, emphasizing its ability to generate detailed point clouds that facilitate structural analysis and inspection. The segmentation of these point clouds into distinct structural components is crucial, particularly for assessing masonry arch bridges, which often lack comprehensive construction documentation. Traditional manual segmentation methods are time-consuming due to the large size of point clouds, prompting the exploration of deep learning (DL) techniques to automate this process. DL approaches, such as PointNet and its successors, have shown promise in point cloud segmentation by eliminating the need for predefined features, although challenges remain regarding computational efficiency and performance on imbalanced datasets.
The paper introduces BridgeNetv2, a novel architecture that integrates a Transformer-based block for local feature extraction from point clouds. By employing a $k$-nearest neighbors (k-NN) approach for attention mechanisms, BridgeNetv2 aims to reduce memory consumption while maintaining accuracy. The architecture is trained on a synthetic dataset of masonry arch bridges and tested on real-world point clouds, addressing the scarcity of suitable data for training segmentation models in civil engineering. The proposed method is designed to be adaptable for various large-scale civil engineering structures, with the architecture’s details and performance evaluations provided in subsequent sections.
Methods
The methodology section describes the architecture of BridgeNetv2, which employs a classic encoder-decoder framework for point cloud processing. The feature encoder consists of four stages that progressively extract and downsample features using random sampling (RS) with a sampling rate of 4, akin to the approach in RandLA-Net (Hu et al., 2020). The input point cloud, initially sized at $(N, 3)$, is transformed into a higher-dimensional representation of $(N, 8)$ through input embedding. Each stage incorporates two feature extraction (FE) layers, similar to those in the original BridgeNet (Jing et al., 2022), which traditionally utilized max pooling for neighbor feature aggregation. However, this method risks losing low-frequency geometric information.
To address this limitation, BridgeNetv2 introduces a Local Geo-Transformer between each FE block, which computes local attention on the point clouds before and after RS. This module enables the aggregation of features from points that are discarded during RS, thereby preserving geometric information from earlier layers. Additionally, a Global Geo-Transformer is implemented to compute self-attention on the lowest resolution features, capturing long-range dependencies. The enhancements in BridgeNetv2 are primarily the inclusion of the Local and Global Geo-Transformers, while the overall architecture remains consistent with that of BridgeNet. For comprehensive details on the original architecture, readers are directed to Jing et al. (2022).
Results
In the results section, the performance of various semantic segmentation methods, including PointNet++, RandLA-Net, BridgeNet, and the proposed BridgeNetv2, is compared. BridgeNetv2 demonstrates superior performance across all evaluated metrics, achieving a 5.5% increase in mean precision ($\text{mean\_Pr}$) and a 9.1% increase in mean Intersection over Union ($\text{mIoU}$) compared to its predecessor, BridgeNet. Notably, the improvements in segmentation accuracy are particularly pronounced in classes other than the spandrel wall, with BridgeNetv2 significantly enhancing the segmentation of the hole class, which is underrepresented in the dataset.
Visual assessments of BridgeNetv2’s predictions reveal a strong alignment with ground truth, characterized by well-defined boundaries between components. However, misclassifications were noted in specific areas, attributed to discrepancies between synthetic training data and real-world geometries, as well as the challenges posed by data imbalance. The authors suggest that refining the simulation process to better address these imbalances could lead to further enhancements in segmentation accuracy.
Discussion
The discussion section of the paper presents the architecture and functionality of the Local and Global Geo-Transformers integrated into BridgeNetv2, a novel deep learning model for semantic segmentation of masonry arch bridge point clouds. The Local Geo-Transformer employs a localized attention mechanism, computing attention based on the features of $k$ nearest neighbors in point clouds, thereby preserving geometric information during subsampling. A rotationally invariant positional embedding, termed Local Geo-Embedding, is introduced to enhance the model’s sensitivity to geometric features while ensuring robustness against rigid transformations. The Global Geo-Transformer, in contrast, computes self-attention across features at a lower resolution, capturing global contextual relationships.
The paper also details the implementation specifics, including the use of the Adam optimizer, data augmentation techniques, and performance metrics such as overall accuracy and mean intersection over union (mIoU). Ablation tests reveal that a consistent $k$ value across layers optimizes performance, while the proposed rotationally invariant positional embedding significantly outperforms traditional methods. Despite a modest increase in parameter size and memory consumption, BridgeNetv2 demonstrates superior performance compared to existing models like PointNet++ and RandLA-Net, indicating its potential for broader applications in civil engineering tasks. Future work aims to extend the model’s applicability to various structural types, necessitating the development of extensive datasets.
