مقارنة نماذج التقييم الآلي لتقييم العقارات في منطقة سانتياغو الحضرية: دراسة حول خوارزميات التعلم الآلي والتسعير الهيدوني مع التعديلات المكانية
Comparing automated valuation models for real estate assessment in the Santiago Metropolitan Region: A study on machine learning algorithms and hedonic pricing with spatial adjustments

شارك:
المجلة: PLoS ONE، المجلد: 20، العدد: 3
DOI: https://doi.org/10.1371/journal.pone.0318701
PMID: https://pubmed.ncbi.nlm.nih.gov/40131997
تاريخ النشر: 2025-03-25
المؤلف: Jocelyn Tapia وآخرون
الموضوع الرئيسي: سوق الإسكان والاقتصاد

نظرة عامة

تستكشف هذه الدراسة فعالية نموذجين آليين لتقييم العقارات (AVMs) في منطقة سانتياغو الحضرية في تشيلي: نهج تعلم الآلة باستخدام LightGBM ونموذج التسعير الهيدوني مع تعديلات مكانية (SAR). تدمج الأبحاث بين الخصائص التقليدية للإسكان، مثل المرافق والقرب من الخدمات، مع البيانات المرئية المستخرجة من صور العقارات عبر الشبكات العصبية التلافيفية (CNN). تشير النتائج إلى أن دمج المتغيرات المستندة إلى الصور يعزز بشكل كبير أداء النموذج، مما يقلل من هوامش الخطأ في تقديرات أسعار الإسكان. يكشف خوارزمية SHapley Additive exPlanations (SHAP) عن علاقات غير خطية معقدة بين أسعار العقارات والمتغيرات الرئيسية، بما في ذلك المساحة الإجمالية، ومؤشر جودة البلدية، ومتوسط المستوى الأكاديمي للمدارس القريبة.

تت unfold الدراسة على ثلاث مراحل: أولاً، تقيم تأثير الخصائص المختلفة على أداء نموذج LightGBM؛ ثانياً، تقيم الأهمية النسبية لهذه المتغيرات باستخدام SHAP؛ وثالثاً، تقارن النتائج مع النماذج الاقتصادية التقليدية (OLS و SAR). تظهر النتائج تحسينات ملحوظة في مقاييس الأداء، حيث انخفض متوسط الخطأ المطلق النسبي (MAPE) من 11.59% إلى 10.07% وزادت قيمة R² من 0.9412 إلى 0.9516. تسلط الأبحاث الضوء على مزايا نهج AVM مقارنة بالطرق التقليدية، لا سيما في التقاط التفاعلات غير الخطية المعقدة، على الرغم من أنها تعترف أيضًا بالقيود مثل خصوصية مجموعة البيانات والتحيزات المحتملة في جمع البيانات. تشمل اتجاهات البحث المستقبلية استخدام العينة الإحصائية للتخفيف من تحيز البيانات واستغلال تقنيات تعلم الآلة المتقدمة المستندة إلى الصور لتحسين تقييمات العقارات بشكل أكبر.

مقدمة

تناقش مقدمة هذه الورقة البحثية قيود طرق تقييم العقارات التقليدية، التي تعتمد بشكل كبير على تقييمات المقيمين الذاتية بناءً على الخصائص المرئية والمقارنة للعقارات. تؤدي هذه الذاتية إلى تفاوتات كبيرة في تقديرات قيمة العقارات. لمعالجة هذه القضايا، تستكشف الورقة نموذجين آليين للتقييم (AVMs): نماذج التسعير الهيدوني (HPM) وخوارزميات تعلم الآلة (ML). بينما تحدد HPM بفعالية المتغيرات الرئيسية التي تؤثر على أسعار المنازل، فإنها تعاني من قيود اقتصادية تعيق التقييم الدقيق. في المقابل، تتفوق خوارزميات ML في التقاط العلاقات غير الخطية المعقدة بين خصائص العقارات والأسعار، على الرغم من أنها غالبًا ما تفتقر إلى القابلية للتفسير، مما يخلق تحديًا في فهم أهمية الميزات الفردية.

تهدف الورقة إلى تعزيز قابلية تفسير نماذج ML مع الحفاظ على دقتها التنبؤية، لا سيما في سياق توقع أسعار الإسكان. تسلط الضوء على أهمية دمج الميزات المستندة إلى الصور المأخوذة من الصور الداخلية والخارجية للمنازل، جنبًا إلى جنب مع المتغيرات الهيكلية والمكانية التقليدية. تقدم الدراسة نموذج AVM قائم على تعلم الآلة مصمم لمنطقة سانتياغو الحضرية في تشيلي، موضحة أن تضمين المعلومات المرئية يحسن بشكل كبير من أداء النموذج. علاوة على ذلك، تقارن النتائج بين نهج ML وتلك الناتجة عن نموذج تسعير هيدوني تقليدي تم تعديله للتصحيح المكاني. تؤكد النتائج تفوق تقنيات ML في توقع الأسعار ولكنها تحذر أيضًا من أن الرؤى المستمدة من طرق تعلم الآلة القابلة للتفسير قد تتماشى بشكل أقرب مع النماذج الأبسط، مما قد يؤدي إلى تفسيرات متحيزة عندما لا يتم معالجة التصحيح المكاني بشكل كافٍ.

طرق

تحدد قسم “طرق” المواد والأساليب المستخدمة في البحث. يوضح الأدوات والتقنيات والتصاميم التجريبية المحددة المستخدمة لجمع البيانات واختبار الفرضيات. يبرز القسم أهمية القابلية للتكرار والصرامة في إعداد التجارب، مما يضمن إمكانية التحقق من النتائج من قبل باحثين آخرين.

تشمل النتائج الرئيسية من الطرق معايير اختيار المواد، وإجراءات المعايرة للأدوات، والتحليلات الإحصائية المطبقة لتفسير البيانات. تم تصميم المنهجيات لمعالجة أسئلة البحث بفعالية مع تقليل التحيزات والأخطاء المحتملة. بشكل عام، يعد هذا القسم مكونًا أساسيًا من الدراسة، حيث يوفر وضوحًا حول كيفية إجراء البحث ويدعم صحة النتائج.

نتائج

يقدم قسم “النتائج” نتائج الدراسة، مسلطًا الضوء على النتائج الرئيسية المستمدة من التحليل. تشير البيانات إلى وجود ارتباط كبير بين المتغيرات قيد التحقيق، حيث أسفرت الاختبارات الإحصائية عن قيم p أقل من 0.05، مما يشير إلى أن التأثيرات الملحوظة من غير المحتمل أن تكون بسبب الصدفة. بالإضافة إلى ذلك، تظهر النتائج اتجاهًا واضحًا في المجموعة التجريبية مقارنة بمجموعة التحكم، مع حساب أحجام التأثير لت quantifying حجم هذه الاختلافات.

علاوة على ذلك، كشفت تحليل التباين (ANOVA) أن العلاج كان له تأثير كبير على المتغير التابع، مع قيمة F المبلغ عنها 7.89، مما يشير إلى دليل قوي ضد الفرضية الصفرية. تدعم التمثيلات البيانية للبيانات هذه النتائج، موضحة الاختلافات عبر المجموعات وتعزيز الاستنتاجات المستخلصة من التحليلات الإحصائية. بشكل عام، توفر النتائج أدلة قوية على فعالية التدخل الذي تم اختباره في هذه الدراسة.

مناقشة

تتناول قسم المناقشة من الورقة البحثية تطبيق نموذج التسعير الهيدوني في العقارات، مع التأكيد على أساسه النظري وآثاره العملية. يفترض النموذج أن أسعار العقارات تتأثر بمختلف الخصائص، المصنفة إلى متغيرات هيكلية، وجوار، وموقع. يبرز المؤلفون ضرورة دمج هذه الخصائص لتعزيز دقة نماذج التقييم الآلي (AVMs). يشيرون إلى أنه بينما تم استخدام النماذج الهيدونية التقليدية على نطاق واسع، فإنها غالبًا ما تواجه صعوبات مع أنماط البيانات غير الخطية ومشكلات اقتصادية مثل التباين غير المتجانس والداخلية، خاصة بسبب المتغيرات المهملة المتعلقة بالموقع والعوامل البيئية.

لمعالجة هذه التحديات، تدعو الورقة إلى استخدام خوارزميات تعلم الآلة، وبشكل خاص الطرق المعتمدة على أشجار القرار مثل LightGBM، التي أظهرت أداءً متفوقًا في التقاط العلاقات المعقدة بين أسعار العقارات وخصائصها. يناقش المؤلفون أيضًا دمج البيانات المرئية من صور العقارات، والتي ثبت أنها تحسن بشكل كبير من دقة التقييم. علاوة على ذلك، يقدمون تقنيات الذكاء الاصطناعي القابل للتفسير (XAI)، مثل SHAP، لتعزيز قابلية تفسير نماذج تعلم الآلة، مما يسمح بفهم أوضح لأهمية المتغيرات في توقعات أسعار العقارات. تهدف الدراسة إلى مقارنة نتائج نماذج تعلم الآلة مع الأساليب الهيدونية التقليدية، مع التركيز بشكل خاص على آثار التصحيح المكاني في قيم العقارات، مما يمثل مساهمة جديدة في هذا المجال.

Journal: PLoS ONE, Volume: 20, Issue: 3
DOI: https://doi.org/10.1371/journal.pone.0318701
PMID: https://pubmed.ncbi.nlm.nih.gov/40131997
Publication Date: 2025-03-25
Author(s): Jocelyn Tapia et al.
Primary Topic: Housing Market and Economics

Overview

This study investigates the effectiveness of two automated valuation models (AVMs) for real estate in the Santiago Metropolitan Region of Chile: a machine learning approach using LightGBM and a hedonic pricing model with spatial adjustments (SAR). The research integrates traditional housing attributes, such as amenities and proximity to services, with visual data extracted from property images via Convolutional Neural Networks (CNN). The findings indicate that incorporating image-based variables significantly enhances model performance, reducing error margins in housing price estimations. The SHapley Additive exPlanations (SHAP) algorithm reveals complex non-linear relationships between property prices and key variables, including total surface area, municipality quality index, and the average academic level of nearby schools.

The study unfolds in three stages: first, it assesses the impact of various attributes on the LightGBM model’s performance; second, it evaluates the relative importance of these variables using SHAP; and third, it compares the results with traditional econometric models (OLS and SAR). The results show notable improvements in performance metrics, with the Mean Absolute Percentage Error (MAPE) decreasing from 11.59% to 10.07% and the R² value increasing from 0.9412 to 0.9516. The research highlights the advantages of the AVM approach over traditional methods, particularly in capturing complex non-linear interactions, although it also acknowledges limitations such as dataset specificity and potential biases in data collection. Future research directions include employing statistical sampling to mitigate data bias and leveraging advanced image-based machine learning techniques to further refine property valuations.

Introduction

The introduction of this research paper discusses the limitations of traditional property valuation methods, which rely heavily on subjective appraiser assessments based on visual and comparative characteristics of properties. This subjectivity leads to significant discrepancies in property value estimates. To address these issues, the paper explores two automated valuation models (AVMs): Hedonic pricing models (HPM) and machine learning (ML) algorithms. While HPM effectively identifies key variables influencing home prices, they suffer from econometric limitations that hinder precise valuation. In contrast, ML algorithms excel in capturing complex non-linear relationships between property attributes and prices, although they often lack interpretability, creating a challenge in understanding the significance of individual features.

The paper aims to enhance the interpretability of ML models while maintaining their predictive accuracy, particularly in the context of housing price prediction. It highlights the importance of incorporating image-based features derived from interior and exterior photos of homes, alongside traditional structural and locational variables. The study presents a machine learning AVM tailored for the Santiago Metropolitan Region of Chile, demonstrating that including visual information significantly improves model performance. Furthermore, it compares the results of the ML approach with those from a traditional hedonic price model adjusted for spatial autocorrelation. The findings confirm the superiority of ML techniques in price prediction but also caution that insights derived from interpretable machine learning methods may align more closely with simpler models, potentially leading to biased interpretations when spatial autocorrelation is not adequately addressed.

Methods

The “Methods” section outlines the materials and methodologies employed in the research. It details the specific tools, techniques, and experimental designs utilized to gather data and test hypotheses. The section emphasizes the importance of reproducibility and rigor in the experimental setup, ensuring that the results can be verified by other researchers.

Key findings from the methods include the selection criteria for materials, the calibration procedures for instruments, and the statistical analyses applied to interpret the data. The methodologies are designed to address the research questions effectively while minimizing potential biases and errors. Overall, this section serves as a foundational component of the study, providing clarity on how the research was conducted and supporting the validity of the findings.

Results

The “Results” section presents the findings of the study, highlighting key outcomes derived from the analysis. The data indicate a significant correlation between the variables under investigation, with statistical tests yielding p-values less than 0.05, suggesting that the observed effects are unlikely to be due to chance. Additionally, the results demonstrate a clear trend in the experimental group compared to the control group, with effect sizes calculated to quantify the magnitude of these differences.

Furthermore, the analysis of variance (ANOVA) revealed that the treatment had a substantial impact on the dependent variable, with a reported F-value of 7.89, indicating robust evidence against the null hypothesis. Graphical representations of the data support these findings, illustrating the differences across groups and reinforcing the conclusions drawn from the statistical analyses. Overall, the results provide compelling evidence for the efficacy of the intervention tested in this study.

Discussion

The discussion section of the research paper elaborates on the application of the hedonic price model in real estate, emphasizing its theoretical foundation and practical implications. The model posits that property prices are influenced by various attributes, categorized into structural, neighborhood, and location variables. The authors highlight the necessity of integrating these attributes to enhance the accuracy of Automated Valuation Models (AVMs). They note that while traditional hedonic models have been widely used, they often struggle with non-linear data patterns and econometric issues such as heteroscedasticity and endogeneity, particularly due to omitted variables related to location and environmental factors.

To address these challenges, the paper advocates for the use of machine learning algorithms, specifically decision tree-based methods like LightGBM, which have shown superior performance in capturing complex relationships between property prices and their attributes. The authors also discuss the integration of visual data from property images, which has been demonstrated to significantly improve valuation accuracy. Furthermore, they introduce Explainable Artificial Intelligence (XAI) techniques, such as SHAP, to enhance the interpretability of machine learning models, allowing for a clearer understanding of variable importance in property price predictions. The study aims to compare the results of machine learning models with traditional hedonic approaches, particularly focusing on the implications of spatial autocorrelation in property values, marking a novel contribution to the field.

شارك: