مسافات جديدة للبيانات المختلطة القادرة على التعامل مع المعلومات الزائدة
New distances for mixed-type data able to cope with redundant information

شارك:
المجلة: AStA Advances in Statistical Analysis
DOI: https://doi.org/10.1007/s10182-026-00565-6
تاريخ النشر: 2026-06-22
المؤلف: Aurea Grané وآخرون
الموضوع الرئيسي: طرق ونماذج إحصائية متقدمة

نظرة عامة

تتناول هذه البحث التكرار الموجود غالبًا في المتغيرات الفئوية التي تم جمعها من خلال الاستطلاعات، مما يمكن أن يؤدي إلى نتائج مضللة عند استخدام تقنيات تعتمد على المسافة مثل القياس متعدد الأبعاد (MDS). يقترح المؤلفون مقاييس عدم التشابه الجديدة للبيانات الفئوية التي تأخذ في الاعتبار الهياكل الارتباطية الكامنة بين المتغيرات، وتدمج هذه مع مقاييس المسافة القوية للبيانات العددية لإنشاء مقياس مرن مناسب لمجموعات البيانات المختلطة. يتم تقييم فعالية هذه المقاييس المقترحة تحت ظروف سلبية متنوعة، بما في ذلك هياكل الارتباط وتلوث القيم الشاذة، ومقارنتها بمسافة غاور الكلاسيكية باستخدام تمثيلات MDS ومصنف الجار الأقرب.

تشير النتائج إلى أن المقاييس المقترحة تدمج بنجاح الهيكل الارتباطي للبيانات، مما يقلل من المشكلات مثل تأثير الحدوة ويعزز اكتشاف القيم الشاذة في تكوينات MDS. يتضمن البحث تطبيقات على مجموعات بيانات من العالم الحقيقي، مما يظهر دقة تصنيف محسنة وتحديد أفضل للقيم الشاذة مقارنة بالطرق الكلاسيكية. يُقترح العمل المستقبلي لتحسين هذه المقاييس المسافة بشكل أكبر واستكشاف قابليتها للتطبيق عبر تقنيات إحصائية متنوعة، مما قد يعزز تكاملها في أطر تحليل البيانات الحالية. بشكل عام، يبرز الدراسة أهمية المنهجية المقترحة في معالجة التكرار في البيانات الفئوية وآثارها على تحليل البيانات بدقة أكبر.

مقدمة

تتناول مقدمة هذه الورقة البحثية التعقيد المتزايد للبيانات التي يواجهها المحللون، خاصة عند التعامل مع مجموعات بيانات غير متجانسة تشمل أنواعًا مختلفة من المتغيرات – الكمية، الثنائية، والفئوية متعددة الفئات. غالبًا ما تفشل الطرق الإحصائية التقليدية في تلبية الافتراضات اللازمة للطبيعية، والتجانس، والاستقلالية، خاصة في وجود القيم الشاذة وهياكل الارتباط. لمواجهة هذه التحديات، يقترح المؤلفون مقاييس مسافة قوية جديدة للبيانات المختلطة تأخذ في الاعتبار الارتباط بين المتغيرات العددية وتهدف إلى معالجة التكرار في البيانات الفئوية.

تنتقد الورقة مسافة غاور، وهي طريقة كلاسيكية للبيانات المختلطة، لقيودها في تجاهل ارتباط المتغيرات وحساسيتها للقيم الشاذة. يبني المؤلفون على الأعمال السابقة التي قدمت مسافات قوية لتطبيقات في التجميع، والتوصيف، والنمذجة التنبؤية. يحددون منهجية تتضمن حساب مقاييس عدم التشابه للبيانات الفئوية، والتي يتم دمجها بعد ذلك مع مسافات ماهالانوبس القوية للبيانات العددية باستخدام القياس المتعلق (RelMS). لا يحافظ هذا النهج فقط على خاصية إقليدس، مما يضمن أن المقياس النهائي هو مقياس مسافة صالح، ولكنه يبسط أيضًا العملية من خلال القضاء على الحاجة لضبط المعلمات. تعد الورقة بإظهار فعالية هذه المقاييس الجديدة من خلال دراسات محاكاة وتطبيقات من العالم الحقيقي، مما يسهم في تحسين تحليل البيانات في سيناريوهات معقدة.

طرق

في هذا القسم، يقترح المؤلفون منهجية جديدة لقياس المسافات في البيانات المختلطة من خلال معالجة الخصائص المميزة للمتغيرات العددية، الاسمية، والترتيبية. تتضمن الطريقة تقسيم مجموعة البيانات بناءً على أنواع المتغيرات وتطبيق مقاييس عدم التشابه المخصصة التي تأخذ في الاعتبار الهياكل الارتباطية الأساسية. استخدمت الدراسات السابقة بشكل أساسي مسافات ماهالانوبس القوية للبيانات العددية، بينما لم يتم فحص الجوانب الفئوية بشكل شامل. يهدف هذا البحث إلى تعزيز معالجة البيانات الفئوية من خلال التمييز بين المتغيرات الاسمية والترتيبية واستخدام مقاييس عدم التشابه المحددة لكل منها، مستندة إلى مصفوفة الارتباط/الارتباط لالتقاط التكرار وتقليد سلوك مسافة ماهالانوبس.

تتوج المنهجية في توليد مصفوفتين تمثلان عدم التشابه الثنائي للبيانات الفئوية، والتي يتم دمجها بعد ذلك مع مصفوفة المسافة للبيانات العددية باستخدام القياس المتعلق (RelMS). يضمن هذا الدمج أن مصفوفة المسافة المشتركة النهائية تلتزم بالمتطلبات الإقليدية، مما يحافظ على خصائص المقياس كما حددها غاور وليجيندر. يتناول القسم أيضًا مقاييس عدم التشابه المقترحة للبيانات الفئوية ويعيد التأكيد على المنهجية العامة لدمج مصفوفات المسافة المختلفة المحسوبة على نفس مجموعة البيانات.

نتائج

في هذا القسم، يتم تقديم نتائج الدراسة، مع التركيز على حساب المصفوفات باستخدام V لكرايمر للأزواج الاسمية وτ لكيندال للأزواج الترتيبية، مع نهج محدد للسيناريو 3 حيث تم التعامل مع جميع المتغيرات الفئوية كاسمية. استخدمت التحليل مقاييس G-Gower وRelMS القوية لتصور البيانات المختلطة من خلال القياس متعدد الأبعاد (MDS)، مما يكشف أن القيم الشاذة تم إخفاؤها بشكل فعال في تكوينات غاور، بينما قدمت RelMS تجميعات أوضح، خاصة عند قيم الارتباط الأعلى ($\rho$). أشارت وجود نمط بارابوليكي في تكوينات غاور إلى تشبع المقياس بسبب المعلومات الزائدة.

أشارت النتائج الإضافية إلى أنه بالنسبة للبيانات الفئوية غير الملوثة مع ارتباط معتدل ($\rho = 0.6$)، أظهرت كل من G-Gower وRelMS هياكل تجميع مشابهة، على الرغم من أن RelMS أظهرت تركيز وحدة أعلى. عززت إدخال مصفوفات الارتباط وضوح التجميع، خاصة مع المتغيرات الثنائية. كشفت إجراء تصنيف باستخدام K-أقرب الجيران (KNN) أن RelMS القوي تفوقت على G-Gower في السيناريوهات ذات الارتباط الأعلى، بينما حققت مسافة المستندة إلى كرايمر أفضل النتائج للبيانات الفئوية. بشكل عام، تؤكد النتائج على أهمية استخدام المسافات القوية لكشف الأنماط في مجموعات البيانات المعقدة التي تنحرف عن الافتراضات التقليدية، مما يسهل تفسير البيانات وتحليلها بشكل أفضل.

مناقشة

في هذا القسم، يناقش المؤلفون تطوير وتقييم مقاييس عدم التشابه المصممة لمجموعات البيانات المختلطة، مع التركيز بشكل خاص على المتغيرات الفئوية. يتم هيكلة مجموعة البيانات كمصفوفة $n \times p$، حيث تكون أول $p_1$ أعمدة عددية، والأعمدة التالية $p_2$ اسمية، والمتبقية $p_3$ متغيرات ترتيبية. يعرف المؤلفون مصفوفات الارتباط $A_\ell$ للأجزاء الاسمية والترتيبية من مجموعة البيانات، باستخدام معاملات الارتباط مثل V لكرايمر للمتغيرات الاسمية وτ لكيندال للمتغيرات الترتيبية. تعتبر هذه المصفوفات ضرورية لاشتقاق مقاييس عدم التشابه، المشار إليها بـ $\delta_2$ و$\delta_3$، والتي تحقق خصائص عدم السلبية، والهوية، والتناظر، على الرغم من أنها قد لا تلبي دائمًا عدم المساواة مثلث.

يتناول القسم بمزيد من التفصيل طريقة تسمى القياس المتعلق (RelMS)، والتي تجمع بين مصفوفات عدم التشابه المختلفة لإنشاء مسافة مقياس موحد تأخذ في الاعتبار أنواع المتغيرات المختلفة. تعتبر هذه الطريقة مفيدة بشكل خاص عند التعامل مع البيانات المختلطة، حيث تسمح بدمج مصادر المعلومات المتنوعة مع تقليل التكرار. يقترح المؤلفون إجراءً منهجيًا لدمج مصفوفات المسافة، مما يضمن أن المسافة المقياسية المشتركة الناتجة تلتزم بخصائص إقليدس. كما يبرزون أهمية معالجة القضايا المحتملة للوجود الإيجابي في مصفوفات الارتباط، مقترحين خطة تقليص إذا لزم الأمر. يختتم القسم بدراسة محاكاة تهدف إلى تقييم قوة وفعالية مقاييس عدم التشابه المقترحة في سيناريوهات متنوعة، بما في ذلك مجموعات البيانات المختلطة والفئوية البحتة.

Journal: AStA Advances in Statistical Analysis
DOI: https://doi.org/10.1007/s10182-026-00565-6
Publication Date: 2026-06-22
Author(s): Aurea Grané et al.
Primary Topic: Advanced Statistical Methods and Models

Overview

This research addresses the redundancy often present in categorical variables collected through surveys, which can lead to misleading results when employing distance-based techniques like Multidimensional Scaling (MDS). The authors propose new dissimilarity measures for categorical data that account for the inherent association structures among variables, integrating these with robust distance metrics for numerical data to create a flexible metric suitable for mixed-type datasets. The effectiveness of these proposed metrics is evaluated under various adverse conditions, including correlation structures and outlier contamination, and compared to the classical Gower distance using MDS representations and a Nearest Neighbor classifier.

The findings indicate that the proposed metrics successfully incorporate the association structure of the data, mitigating issues such as the horseshoe effect and enhancing outlier detection in MDS configurations. The research includes applications to real-world datasets, demonstrating improved classification accuracy and better identification of outliers compared to classical methods. Future work is suggested to refine these distance metrics further and explore their applicability across various statistical techniques, potentially enhancing their integration into existing data analysis frameworks. Overall, the study highlights the relevance of the proposed methodology in addressing redundancy in categorical data and its implications for more accurate data analysis.

Introduction

The introduction of this research paper addresses the increasing complexity of data encountered by analysts, particularly when dealing with heterogeneous datasets that include various types of variables—quantitative, binary, and multiclass categorical. Traditional statistical methods often fail to meet the necessary assumptions of normality, homogeneity, and independence, especially in the presence of outliers and correlation structures. To tackle these challenges, the authors propose new robust distance measures for mixed-type data that account for the correlation among numerical variables and aim to address redundancy in categorical data.

The paper critiques Gower’s distance, a classical method for mixed-type data, for its limitations in ignoring variable associations and its sensitivity to outliers. The authors build upon previous works that introduced robust distances for applications in clustering, profiling, and predictive modeling. They outline a methodology that involves computing dissimilarity measures for categorical data, which are then integrated with robust Mahalanobis distances for numerical data using Related Metric Scaling (RelMS). This approach not only maintains the Euclidean property, ensuring the final metric is a valid distance measure, but also simplifies the process by eliminating the need for tuning parameters. The paper promises to demonstrate the effectiveness of these new measures through simulation studies and real-world applications, ultimately contributing to improved data analysis in complex scenarios.

Methods

In this section, the authors propose a novel methodology for measuring distances in mixed-type data by addressing the distinct characteristics of numerical, nominal, and ordinal variables. The approach involves partitioning the dataset based on variable types and applying tailored dissimilarity measures that account for the underlying correlation structures. Previous studies primarily utilized robust Mahalanobis distances for numerical data, while the categorical aspects were not thoroughly examined. This research aims to enhance the treatment of categorical data by differentiating between nominal and ordinal variables and employing specific dissimilarity measures for each, informed by an association/correlation matrix to capture redundancy and mimic the behavior of Mahalanobis distance.

The methodology culminates in the generation of two matrices representing pairwise dissimilarities for categorical data, which are then integrated with the distance matrix for numerical data using Related Metric Scaling (RelMS). This integration ensures that the final joint distance matrix adheres to the Euclidean requirement, thereby maintaining the metric properties as established by Gower and Legendre. The section also outlines two proposed dissimilarity measures for categorical data and reiterates the general methodology for combining various distance matrices computed on the same dataset.

Results

In this section, the results of the study are presented, focusing on the computation of matrices using Cramér’s V for nominal pairs and Kendall’s τ for ordinal pairs, with a specific approach for Scenario 3 where all categorical variables were treated as nominal. The analysis employed robust G-Gower and RelMS metrics to visualize mixed-type data through Multi-Dimensional Scaling (MDS), revealing that outliers were effectively masked in Gower’s configurations, while RelMS provided clearer groupings, particularly at higher correlation values ($\rho$). The presence of a parabolic pattern in Gower’s configurations suggested metric saturation due to redundant information.

Further results indicated that for uncontaminated categorical data with moderate association ($\rho = 0.6$), both G-Gower and RelMS demonstrated similar grouping structures, although RelMS exhibited higher unit concentration. The introduction of association matrices enhanced grouping clarity, especially with binary variables. A classification procedure using K-Nearest Neighbors (KNN) revealed that robust RelMS outperformed G-Gower in scenarios with higher correlation, while Cramér-based distance yielded the best results for categorical data. Overall, the findings underscore the importance of employing robust distances to uncover patterns in complex datasets that deviate from traditional assumptions, thereby facilitating better data interpretation and analysis.

Discussion

In this section, the authors discuss the development and evaluation of dissimilarity measures tailored for mixed-type datasets, specifically focusing on categorical variables. The dataset is structured as an $n \times p$ matrix, where the first $p_1$ columns are numerical, the next $p_2$ are nominal, and the remaining $p_3$ are ordinal variables. The authors define association matrices $A_\ell$ for the nominal and ordinal parts of the dataset, utilizing association coefficients such as Cramér’s V for nominal variables and Kendall’s τ for ordinal variables. These matrices are crucial for deriving dissimilarity measures, denoted as $\delta_2$ and $\delta_3$, which fulfill properties of non-negativity, identity, and symmetry, although they may not always satisfy the triangle inequality.

The section further elaborates on a method called Related Metric Scaling (RelMS), which combines various dissimilarity matrices to create a unified metric distance that accounts for different variable types. This approach is particularly beneficial when dealing with mixed-type data, as it allows for the integration of diverse information sources while minimizing redundancy. The authors propose a systematic procedure for combining distance matrices, ensuring that the resulting joint metric distance adheres to Euclidean properties. They also highlight the importance of addressing potential issues of positive definiteness in the association matrices, suggesting a shrinkage scheme if necessary. The section concludes with a simulation study aimed at assessing the robustness and effectiveness of the proposed dissimilarity measures in various scenarios, including mixed-type and purely nominal datasets.

شارك: