استنتاج وظيفة حضرية بدون عينة باستخدام صور عرض الشارع من خلال تحفيز نموذج رؤية-لغة مدرب مسبقًا
Zero-shot urban function inference with street view images through prompting a pretrained vision-language model

شارك:
المجلة: International Journal of Geographical Information Systems، المجلد: 38، العدد: 7
DOI: https://doi.org/10.1080/13658816.2024.2347322
تاريخ النشر: 2024-05-22
المؤلف: Weiming Huang وآخرون
الموضوع الرئيسي: تنقل البشر والتحليل القائم على الموقع

نظرة عامة

تقدم ورقة البحث UrbanCLIP، وهو إطار عمل جديد مصمم للاستفادة من نموذج اللغة-الصورة المدرب مسبقًا CLIP من أجل الاستدلال على الوظائف الحضرية بدون تدريب مسبق باستخدام صور عرض الشارع (SVIs). يعالج هذا الإطار تحديات كبيرة في استدلال الوظائف الحضرية، مثل الاعتماد على البيانات المعلّمة والموارد الحاسوبية الكبيرة. يدمج UrbanCLIP تصنيفًا حضريًا وقوالب تحفيزية متنوعة لربط فئات الوظائف الحضرية المجردة بأنواع الكائنات الحضرية الملموسة، بينما يقلل أيضًا من التداخل من العناصر الخارجية في SVIs، مثل الأشجار والمركبات.

تظهر التجارب الواسعة أن UrbanCLIP يتفوق بشكل كبير على عدة نماذج إشرافية تنافسية، بما في ذلك نموذج ResNet المعدل، خاصة في سيناريوهات نقل البيانات بين المدن. تشير النتائج إلى أن أداء UrbanCLIP بدون تدريب مسبق يتفوق بشكل ملحوظ على أداء نموذج CLIP القياسي. بشكل عام، لا يبسط الإطار عملية استدلال الوظائف الحضرية فحسب، بل يقلل أيضًا من الحاجة إلى عينات معلّمة وموارد حاسوبية، مما يظهر إمكانيات النماذج الأساسية في التطبيقات الجغرافية.

مقدمة

تؤكد مقدمة ورقة البحث هذه على أهمية فهم التوزيع المكاني للوظائف الحضرية من أجل إدارة حضرية فعالة وتطوير مستدام. تبرز تعقيد البيئات الحضرية، التي تتكون من مناطق وظيفية متنوعة مثل المناطق السكنية والتجارية والصناعية. يتم دعم الطرق التقليدية للتعرف على الوظائف الحضرية، بما في ذلك الاستطلاعات الميدانية والصور الجوية، بشكل متزايد بتقنيات استخراج البيانات التي تستخدم مصادر بيانات استشعار حضرية متنوعة، وخاصة صور عرض الشارع (SVIs). تُعرف SVIs بقدرتها على توفير معلومات بصرية مفصلة يمكن أن تميز بفعالية بين الوظائف الحضرية المختلفة، على الرغم من أن التحديات لا تزال قائمة بسبب الحاجة إلى مجموعات بيانات موسومة كبيرة لتدريب نماذج التعلم العميق.

لمعالجة هذه القيود، تقدم الورقة UrbanCLIP، وهو إطار عمل جديد مصمم للاستفادة من قدرات نموذج CLIP (التدريب المسبق على اللغة والصورة المتباينة) لاستدلال الوظائف الحضرية بدون تدريب مسبق. يهدف UrbanCLIP إلى تبسيط العملية من خلال عدم الحاجة إلى عينات موسومة أو إعادة تدريب النموذج، مما يمكّن من استدلال دقيق للوظائف الحضرية بطريقة بدون تدريب مسبق. يدمج الإطار تصنيفًا حضريًا لربط فئات الوظائف الحضرية المجردة بأنواع الكائنات الحضرية الملموسة ويستخدم قوالب تحفيزية لتقليل الضوضاء من العناصر البصرية الشائعة في SVIs. تظهر النتائج التجريبية أن UrbanCLIP يتفوق على عدة نماذج أساسية تحت إشراف في استدلال الوظائف الحضرية الأساسية والمتعددة عبر مدن مختلفة، مما يظهر قوته وقابليته للنقل في السياقات الحضرية.

طرق البحث

في هذه الدراسة، يقيم المؤلفون أداء نموذج UrbanCLIP بدون تدريب مسبق عبر ثلاثة إعدادات تجريبية متميزة. يركز الإعداد الأول على تصنيف الوظائف الأساسية، حيث يهدف النموذج إلى تحديد الوظيفة الرئيسية الممثلة في كل صورة عرض شارع (SVI). يقيم الإعداد الثاني قدرة UrbanCLIP على تصنيف وظائف متعددة، بما في ذلك الوظائف الأساسية والثانوية، ضمن SVIs الفردية. أخيرًا، يتضمن الإعداد الثالث نقل البيانات بين المدن، حيث يقارن المؤلفون أداء UrbanCLIP والنماذج الأساسية في بيئتين حضريتين إضافيتين، سنغافورة ولندن، لتحليل قدرات النقل الخاصة بهم. تهدف هذه التقييمات الشاملة إلى تسليط الضوء على قوة UrbanCLIP وقدرته على التكيف في سياقات حضرية متنوعة.

النتائج

يقدم قسم النتائج في ورقة البحث نتائج حول فعالية UrbanCLIP في تصنيف الوظائف الحضرية على مستوى المشهد، مع التركيز على أدائه في تصنيفات الوظائف الأساسية والمتعددة. يتفوق UrbanCLIP بشكل كبير على النماذج الأساسية، بما في ذلك الطرق بدون تدريب مسبق، خاصة في استدلال الوظائف الأساسية، حيث يحقق درجات F1 موزونة تبلغ 0.82 للسكن، 0.72 للتجارة، و0.71 للوظائف الخارجية والطبيعية. ومع ذلك، يواجه صعوبة مع وظائف الفنادق (F1: 0.46) والوظائف المدنية (F1: 0.34)، وغالبًا ما يصنفها بشكل خاطئ بسبب التشابهات البصرية مع فئات أخرى. توضح مصفوفة الالتباس هذه التحديات، مما يبرز نقاط القوة والضعف في النموذج في التعرف على الوظائف الحضرية.

في سياق تصنيف الوظائف المتعددة، يحافظ UrbanCLIP على ميزة تنافسية، خاصة مع وجود عينات موسومة محدودة، على الرغم من أن أدائه يتراجع مقارنة باستدلال الوظائف الأساسية. تشير الورقة إلى أن الخبراء البشريين يتفوقون على UrbanCLIP في هذا الإعداد، مما يقترح طريقًا محتملاً للعمل المستقبلي الذي يجمع بين UrbanCLIP واكتشاف الكائنات لتعزيز التعرف على الوظائف المتعددة. بالإضافة إلى ذلك، تكشف تجارب نقل النموذج بين المدن عن قوة UrbanCLIP، حيث يتفوق في الاستدلال بدون تدريب مسبق عبر مدن مثل سنغافورة ولندن، متفوقًا على النماذج الأساسية المدربة على بيانات شنتشن. تؤكد هذه الأداءات على قابلية تعميم النموذج، بينما تشير أيضًا إلى أن ضبط نماذج الرؤية العميقة قد لا يؤدي إلى فوائد في البيئات الحضرية المتنوعة. بشكل عام، يظهر UrbanCLIP إمكانيات كبيرة لاستدلال الوظائف الحضرية بطريقة بدون تدريب مسبق، على الرغم من بعض القيود مقارنة بالخبرة البشرية.

المناقشة

تسلط قسم المناقشة في ورقة البحث الضوء على الاهتمام المتزايد في استخدام صور عرض الشارع (SVIs) لاستدلال الوظائف الحضرية، مع التأكيد على التقدم الذي تم إحرازه من خلال نماذج التعلم الآلي المختلفة. تشمل المساهمات البارزة تكيف الشبكات العصبية التلافيفية (CNNs) لتصنيف استخدام الأراضي على مستوى المباني وتحويل SVIs إلى أوصاف نصية لتصنيف المشاهد الحضرية. على الرغم من هذه التقدمات، لا تزال التحديات قائمة، خاصة في الحصول على بيانات موسومة كافية لتدريب النماذج، حيث غالبًا ما تعاني بيانات الحقيقة الأرضية من مصادر مثل OpenStreetMap (OSM) من عدم الدقة وعدم الاكتمال. يقترح المؤلفون إطار عمل UrbanCLIP لمعالجة هذه القيود، مستفيدين من التدريب المسبق على اللغة والصورة وهندسة التحفيز لتعزيز استدلال الوظائف الحضرية.

يدمج UrbanCLIP تصنيفًا حضريًا يربط بين عشرة فئات وظيفية حضرية مفصلة و354 نوعًا ملموسًا من الكائنات الحضرية (UOTs)، مما يسهل فهم المشاهد الحضرية بواسطة نموذج CLIP. بالإضافة إلى ذلك، يستخدم الإطار قوالب تحفيزية لوظائف حضرية مصممة لتقليل التداخل من المعلومات البصرية غير ذات الصلة في SVIs، مما يسمح باستدلال أكثر دقة للوظائف الحضرية بدون تدريب مسبق. يجادل المؤلفون بأن نهجهم يمثل مساهمة جديدة في كل من علم نظم المعلومات الجغرافية وهندسة التحفيز، حيث يبرز أهمية المعرفة بالمجال في تطوير استراتيجيات تحفيزية فعالة لنماذج اللغة-الصورة المدربة مسبقًا. بشكل عام، يهدف UrbanCLIP إلى تحسين دقة وكفاءة تصنيف الوظائف الحضرية، خاصة في البيئات الحضرية المعقدة مثل شنتشن، الصين.

Journal: International Journal of Geographical Information Systems, Volume: 38, Issue: 7
DOI: https://doi.org/10.1080/13658816.2024.2347322
Publication Date: 2024-05-22
Author(s): Weiming Huang et al.
Primary Topic: Human Mobility and Location-Based Analysis

Overview

The research paper introduces UrbanCLIP, a novel prompting framework designed to leverage the pretrained vision-language model CLIP for zero-shot inference of urban functions using street view images (SVIs). This framework addresses significant challenges in urban function inference, such as the dependence on labeled data and extensive computational resources. UrbanCLIP integrates an urban taxonomy and various prompt templates to connect abstract urban function categories with concrete urban object types, while also reducing interference from extraneous elements in SVIs, like trees and vehicles.

Extensive experiments demonstrate that UrbanCLIP significantly outperforms several competitive supervised models, including a fine-tuned ResNet, particularly in cross-city transfer scenarios. The results indicate that UrbanCLIP’s zero-shot performance is markedly superior to that of the standard CLIP model. Overall, the framework not only simplifies the process of urban function inference but also reduces the need for labeled samples and computational resources, showcasing the potential of foundation models in geospatial applications.

Introduction

The introduction of this research paper emphasizes the significance of understanding the spatial distribution of urban functions for effective urban management and sustainable development. It highlights the complexity of urban environments, which consist of diverse functional areas such as residential, commercial, and industrial zones. Traditional methods for urban function recognition, including field surveys and aerial imagery, are increasingly supplemented by data mining techniques utilizing various urban sensing data sources, particularly Street View Images (SVIs). SVIs are noted for their ability to provide detailed visual information that can effectively distinguish between different urban functions, although challenges remain due to the need for extensive labeled datasets for training deep learning models.

To address these limitations, the paper introduces UrbanCLIP, a novel prompting framework designed to leverage the capabilities of the CLIP (Contrastive Language-Image Pretraining) model for zero-shot urban function inference. UrbanCLIP aims to simplify the process by requiring no labeled samples or model retraining, thereby enabling fine-grained urban function inference in a zero-shot manner. The framework incorporates an urban taxonomy to connect abstract urban function categories with concrete urban object types and utilizes prompt templates to reduce noise from common visual elements in SVIs. Experimental results demonstrate that UrbanCLIP outperforms several supervised baseline models in inferring primary and multiple urban functions across different cities, showcasing its robustness and transferability in urban contexts.

Methods

In this study, the authors assess the performance of the zero-shot UrbanCLIP model across three distinct experimental settings. The first setting focuses on primary function classification, where the model aims to identify the main function represented in each street view image (SVI). The second setting evaluates UrbanCLIP’s ability to classify multiple functions, including both primary and secondary functions, within individual SVIs. Lastly, the third setting involves cross-city transfer, wherein the authors compare the performance of UrbanCLIP and baseline models in two additional urban environments, Singapore and London, to analyze their transfer capabilities. This comprehensive evaluation aims to highlight the robustness and adaptability of UrbanCLIP in diverse urban contexts.

Results

The results section of the research paper presents findings on the effectiveness of UrbanCLIP for urban function classification at the scene level, emphasizing its performance in both primary and multiple function classifications. UrbanCLIP significantly outperforms baseline models, including zero-shot methods, particularly in primary function inference, where it achieves weighted F1 scores of 0.82 for residential, 0.72 for commercial, and 0.71 for outdoors and natural functions. However, it struggles with hotel (F1: 0.46) and civic functions (F1: 0.34), often misclassifying them due to visual similarities with other categories. The confusion matrix further illustrates these challenges, highlighting the model’s strengths and weaknesses in recognizing urban functions.

In the context of multiple function classification, UrbanCLIP maintains a competitive edge, particularly with limited labeled samples, although its performance declines compared to primary function inference. The paper notes that human experts outperform UrbanCLIP in this setting, suggesting a potential avenue for future work that combines UrbanCLIP with object detection to enhance multi-function recognition. Additionally, the cross-city model transfer experiments reveal UrbanCLIP’s robustness, as it excels in zero-shot inference across cities like Singapore and London, outperforming supervised baselines trained on Shenzhen data. This performance underscores the model’s generalizability, while also indicating that fine-tuning deep vision models may not yield benefits in diverse urban environments. Overall, UrbanCLIP demonstrates significant potential for urban function inference in a zero-shot manner, despite some limitations compared to human expertise.

Discussion

The discussion section of the research paper highlights the growing interest in utilizing Street View Images (SVIs) for urban function inference, emphasizing the advancements made through various machine learning models. Notable contributions include the adaptation of convolutional neural networks (CNNs) for building-level land use classification and the transformation of SVIs into textual descriptions for urban scene classification. Despite these advancements, challenges persist, particularly in acquiring sufficient labeled data for training models, as ground truth data from sources like OpenStreetMap (OSM) often suffer from inaccuracies and incompleteness. The authors propose the UrbanCLIP framework to address these limitations, leveraging vision-language pre-training and prompt engineering to enhance urban function inference.

UrbanCLIP integrates an urban taxonomy that maps ten detailed urban function categories to 354 concrete Urban Object Types (UOTs), facilitating better comprehension of urban scenes by the CLIP model. Additionally, the framework employs urban function prompt templates designed to mitigate interference from irrelevant visual information in SVIs, allowing for more accurate zero-shot inference of urban functions. The authors argue that their approach represents a novel contribution to both GIScience and prompt engineering, as it emphasizes the importance of domain knowledge in developing effective prompting strategies for pretrained vision-language models. Overall, UrbanCLIP aims to improve the accuracy and efficiency of urban function classification, particularly in complex urban environments like Shenzhen, China.

شارك: