DOI: https://doi.org/10.3389/fdgth.2025.1576290
PMID: https://pubmed.ncbi.nlm.nih.gov/40343213
تاريخ النشر: 2025-04-24
المؤلف: Mikel Hernandez وآخرون
الموضوع الرئيسي: التقنيات التي تحافظ على الخصوصية في البيانات
نظرة عامة
تقدم الورقة إطار تقييم شامل لتوليد البيانات الجدولية الاصطناعية، لا سيما في سياق الرعاية الصحية، حيث تعتبر خصوصية المرضى أمرًا بالغ الأهمية. تتناول التحدي المستمر في تحقيق التوازن بين الدقة، والفائدة، والخصوصية في مجموعات البيانات الطبية المعقدة. يجمع الإطار بين مقاييس مختلفة وتدابير مخاطر الخصوصية في ثلاث فئات: الدقة، والفائدة، والخصوصية، بما في ذلك مقياس لتبادل الدقة والفائدة. قام المؤلفون بتطبيق هذا الإطار على ثلاث مجموعات بيانات طبية مفتوحة المصدر، حيث تم تقييم البيانات الاصطناعية التي تم إنشاؤها بواسطة خمسة نماذج توليد مختلفة، سواء مع أو بدون خصوصية تفاضلية.
تشير النتائج إلى أن النماذج التوليدية الأبسط تميل إلى تحقيق دقة وفائدة أعلى، بينما ترتبط النماذج الأكثر تعقيدًا بمخاطر خصوصية أقل. يعزز دمج الخصوصية التفاضلية الحفاظ على الخصوصية ولكنه غالبًا ما يأتي على حساب تقليل الدقة والفائدة. تعترف الدراسة بالقيود، مثل غياب مقاييس التقييم لوقت تدريب النموذج واستخدام الموارد، والاعتماد على المعلمات الافتراضية، وتقييم الخصوصية التفاضلية بميزانية خصوصية واحدة. يقترح المؤلفون اتجاهات البحث المستقبلية، بما في ذلك تحسين المعلمات، واستكشاف آليات خصوصية بديلة، وتطبيقات أوسع للإطار عبر مجموعات بيانات ومجالات متنوعة، مع التأكيد على الحاجة للتعاون مع الأطباء لتقييم الفائدة السريرية. تمهد هذه العمل الطريق لتقدم مشاركة البيانات التي تحافظ على الخصوصية في الرعاية الصحية من خلال تحسين تقييم البيانات الجدولية الاصطناعية.
مقدمة
تناقش مقدمة هذه الورقة البحثية التقدم في تقنيات الصحة الرقمية التي تعزز التشخيص والعلاج الشخصي من خلال تحليل البيانات المتقدم. ومع ذلك، فإن الاستخدام الفعال لهذه التقنيات يعيقه القلق بشأن الخصوصية المرتبطة بالبيانات الحساسة، مما يؤدي إلى تنظيمات صارمة مثل GDPR وHIPAA التي تحد من مشاركة البيانات. لمعالجة هذه التحديات، ظهرت توليد البيانات الجدولية الاصطناعية (STDG) كتقنية تعزز الخصوصية التي تنشئ بيانات واقعية وغير قابلة للتعريف مع الحفاظ على الخصائص الإحصائية لمجموعات البيانات الأصلية. على الرغم من إمكانياتها، تواجه تطبيقات STDG في الرعاية الصحية تحديات تتعلق بالحفاظ على دقة البيانات، والفائدة، والخصوصية، مما يستلزم إطار تقييم قوي.
تحدد الورقة الفجوات في أطر تقييم البيانات الاصطناعية الحالية، التي غالبًا ما تفتقر إلى توافق حول المقاييس وتدابير مخاطر الخصوصية، مما يجعل التوحيد القياسي صعبًا. تقترح إطار تقييم شامل يقيم بدقة البيانات الجدولية الاصطناعية التي تم إنشاؤها بواسطة نماذج STDG المختلفة، مع دمج مجموعة الحد الأدنى من المقاييس التي تركز على الدقة، والفائدة، والخصوصية. يتضمن هذا الإطار أيضًا مقياسًا لتحديد تبادل الدقة والفائدة تحت قيود الخصوصية المختلفة. يتم إثبات فعالية الإطار المقترح من خلال تطبيقه على ثلاث مجموعات بيانات طبية مفتوحة المصدر، حيث يتم مقارنة البيانات الاصطناعية التي تم إنشاؤها مع وبدون خصوصية تفاضلية (DP). تشير النتائج إلى أن الإطار يمكنه تقييم ومقارنة البيانات الاصطناعية بفعالية، مما يوفر رؤى حول تأثير دمج DP على خصائص البيانات ويوجه القرارات المستنيرة لحلول الرعاية الصحية المعتمدة على البيانات.
الطرق
تصف الإجراءات التجريبية الموضحة في هذا القسم تطبيق إطار تقييم البيانات الجدولية الاصطناعية. في البداية، يتم تقسيم مجموعة بيانات حقيقية عشوائيًا إلى مجموعات تدريب (80%) واختبار (20%)، مع ضمان القابلية للتكرار من خلال استخدام بذور عشوائية ثابتة. يتم الحفاظ على هذا الانقسام الفردي طوال التقييم لتقليل التكاليف الحاسوبية المرتبطة بإعادة تدريب نماذج توليد البيانات الاصطناعية. يتم استخدام منهجية TRTR بعد ذلك باستخدام نماذج التصنيف والانحدار، مما يحدد نتائج أساسية لتقييم فائدة البيانات الجدولية الاصطناعية.
بعد التقييم الأساسي، يتم إنشاء مجموعة بيانات اصطناعية تتطابق مع حجم مجموعة التدريب باستخدام نموذج توليد البيانات الجدولية الاصطناعية (STDG). ثم يتم تطبيق إطار التقييم على هذه المجموعة الاصطناعية عبر 10 طيات، محاكيًا نهج التحقق المتقاطع. يتم تجميع النتائج من كل طية لحساب أفضل وأسوأ ومتوسط القيم لمقاييس مختلفة وتدابير مخاطر الخصوصية. يتم تنفيذ العملية بالكامل في بايثون، باستخدام منصة MLflow مفتوحة المصدر لتتبع النماذج وتنظيم النتائج.
النتائج
ت outlines section النتائج النتائج من تطبيق إطار تقييم البيانات الجدولية الاصطناعية المقترح، كما هو موضح في القسم 2.2. يتم تلخيص مقاييس التقييم وتدابير مخاطر الخصوصية لثلاث مجموعات بيانات، مع التأكيد على تأثير الخصوصية التفاضلية (DP) على نماذج توليد البيانات الجدولية الاصطناعية (STDG). يتم تقديم مقارنة مفصلة لأداء نماذج STDG المختلفة، مع التركيز على تحديد النموذج الأفضل أداءً.
تقدم الجدول 12 ملخصًا شاملاً للنتائج عبر جميع فئات التقييم ومجموعات البيانات، مما يشير إلى أفضل وأسوأ نماذج STDG. يتم تصنيف تأثير DP على أنه إيجابي أو سلبي أو محايد بناءً على تأثيره على غالبية المقاييس أو تدابير مخاطر الخصوصية، مع تحديد الأهمية عند عتبة $p < 0.05$. يسمح هذا التصنيف بفهم دقيق لكيفية تأثير دمج DP على أداء النموذج ونتائج الخصوصية. تتوفر مزيد من التفاصيل في المواد التكميلية.
المناقشة
يؤكد إطار تقييم البيانات الجدولية الاصطناعية المقترح على أربع فئات تقييم رئيسية: الدقة، والفائدة، والخصوصية، وتبادل الدقة والفائدة. هذه الفئات حاسمة لتقييم مدى تقليد البيانات الاصطناعية للبيانات الحقيقية (الدقة)، وقابليتها للتطبيق في السيناريوهات العملية (الفائدة)، وفعاليتها في حماية المعلومات الحساسة (الخصوصية). يستخدم الإطار مقاييس مختلفة، بما في ذلك مسافة هيلينجر للدقة، التي تقيس التشابه بين توزيعات الاحتمالات للبيانات الحقيقية والاصطناعية، واختلاف الارتباط الثنائي (PCD)، الذي يحدد الحفاظ على الارتباطات بين المتغيرات. بالإضافة إلى ذلك، يتم استخدام مخطط العمق مقابل العمق (DD-Plot) ومساحة تحت منحنى التشغيل الاستقبالي (AUC-ROC) لتقييم التشابه التوزيعي المتعدد المتغيرات والتمييز بين العينات الحقيقية والاصطناعية، على التوالي.
لتقييم الفائدة، يقارن الإطار أداء نماذج التعلم الآلي باستخدام البيانات الاصطناعية مقابل البيانات الحقيقية، باستخدام مقاييس مثل الدقة ومتوسط الخطأ المطلق. يتم التقاط تبادل الدقة والفائدة بواسطة مقياس G، الذي يوازن بين الدقة والفائدة تحت قيود الخصوصية المتغيرة. يتم دمج تدابير مخاطر الخصوصية، بما في ذلك التمييز، والربط، ومخاطر استنتاج العضوية، أيضًا لتقييم فعالية البيانات الاصطناعية في حماية المعلومات الحساسة. يتم التحقق من تطبيق الإطار باستخدام ثلاث مجموعات بيانات مرتبطة بالصحة، مما يوضح قدرته على تحليل تأثير الخصوصية التفاضلية على جودة البيانات الاصطناعية وتحديد نماذج توليد البيانات الاصطناعية الأكثر ملاءمة للتطبيقات الطبية.
القيود
تسلط قسم القيود في الورقة الضوء على عدة مجالات للتحسين في إطار التقييم المقترح لنماذج توليد البيانات الجدولية الاصطناعية (STDG)، لا سيما في سياق الدقة، والفائدة، والخصوصية. بينما تعتبر هذه الفئات الثلاث ضرورية، يفتقر الإطار إلى مقاييس تقييم إضافية تتعلق بكفاءة النموذج، مثل وقت التدريب، واستهلاك الموارد، وقدرة التعميم، وقابلية التفسير. يجب أن تهدف الأبحاث المستقبلية إلى دمج هذه المقاييس لتوفير تحليل أكثر شمولاً للتكاليف الحاسوبية واستكشاف قابلية تفسير نماذج STDG، مما قد يعزز اختيار النموذج لمختلف التطبيقات.
علاوة على ذلك، كانت المنهجية المستخدمة، التي تضمنت توليد وتقييم 10 طيات للتحقق، تستغرق وقتًا طويلاً. يقترح المؤلفون تحسين حساب مقاييس التقييم وتدابير مخاطر الخصوصية، بالإضافة إلى استخدام تقسيمات تدريب واختبار متعددة لضمان التمثيل والاتساق. قد لا تعكس الاعتماد على المعلمات الافتراضية لنماذج STDG أدائها الأمثل؛ وبالتالي، يجب النظر في ضبط المعلمات في الدراسات المستقبلية. بالإضافة إلى ذلك، تم تقييم تأثير الخصوصية التفاضلية (DP) على أداء النموذج باستخدام ميزانية خصوصية واحدة، مما يشير إلى الحاجة لاستكشاف قيود الخصوصية المتغيرة وآليات الخصوصية البديلة. يشير المؤلفون أيضًا إلى نماذج معينة تحتاج إلى تحسين، مثل DP-TabDif، التي أدت أداءً ضعيفًا في الدقة والفائدة على الرغم من نتائج الخصوصية القوية. أخيرًا، يمكن توسيع تطبيق الإطار ليشمل مجموعات بيانات ومجالات متنوعة، بما في ذلك البيئات الفيدرالية، لتقييم مرونته وفائدته في العالم الحقيقي بالتعاون مع الخبراء السريريين.
DOI: https://doi.org/10.3389/fdgth.2025.1576290
PMID: https://pubmed.ncbi.nlm.nih.gov/40343213
Publication Date: 2025-04-24
Author(s): Mikel Hernandez et al.
Primary Topic: Privacy-Preserving Technologies in Data
Overview
The paper presents a comprehensive evaluation framework for synthetic tabular data generation, particularly in the context of healthcare, where patient privacy is paramount. It addresses the ongoing challenge of balancing fidelity, utility, and privacy in complex medical datasets. The framework consolidates various metrics and privacy risk measures into three categories: fidelity, utility, and privacy, including a fidelity-utility tradeoff metric. The authors applied this framework to three open-source medical datasets, assessing synthetic data generated by five different generative models, both with and without differential privacy.
Findings indicate that simpler generative models tend to achieve higher fidelity and utility, while more complex models are associated with lower privacy risks. The incorporation of differential privacy enhances privacy preservation but often comes at the cost of reduced fidelity and utility. The study acknowledges limitations, such as the absence of evaluation metrics for model training time and resource usage, reliance on default parameters, and the evaluation of differential privacy with a single privacy budget. The authors suggest future research directions, including parameter optimization, exploration of alternative privacy mechanisms, and broader applications of the framework across various datasets and domains, emphasizing the need for collaboration with clinicians to assess clinical utility. This work lays the groundwork for advancing privacy-preserving data sharing in healthcare through improved evaluation of synthetic tabular data.
Introduction
The introduction of this research paper discusses the advancements in digital health technologies that enhance diagnostics and personalized treatment through sophisticated data analysis. However, the effective utilization of these technologies is hindered by privacy concerns associated with sensitive data, leading to strict regulations like GDPR and HIPAA that limit data sharing. To address these challenges, synthetic tabular data generation (STDG) has emerged as a privacy-enhancing technology that creates realistic, non-identifiable data while preserving the statistical properties of the original datasets. Despite its potential, the application of STDG in healthcare faces challenges related to maintaining data fidelity, utility, and privacy, necessitating a robust evaluation framework.
The paper identifies gaps in existing synthetic data evaluation frameworks, which often lack consensus on metrics and privacy risk measures, making standardization difficult. It proposes a comprehensive evaluation framework that rigorously assesses synthetic tabular data generated by various STDG models, integrating a minimal set of metrics focused on fidelity, utility, and privacy. This framework also incorporates a metric for quantifying the tradeoff between fidelity and utility under different privacy constraints. The effectiveness of the proposed framework is demonstrated through its application to three open-source medical datasets, comparing synthetic data generated with and without differential privacy (DP). The findings indicate that the framework can effectively evaluate and compare synthetic data, providing insights into the impact of incorporating DP on data characteristics and guiding informed decisions for data-driven healthcare solutions.
Methods
The experimental procedure outlined in this section describes the application of a synthetic tabular data evaluation framework. Initially, a real dataset is randomly divided into training (80%) and test (20%) sets, ensuring reproducibility through a fixed random seed. This single split is maintained throughout the evaluation to minimize computational costs associated with retraining synthetic data generation models. The TRTR methodology is subsequently employed using classification and regression models, establishing baseline results for assessing the utility of the synthetic tabular data.
Following the baseline evaluation, a synthetic dataset matching the size of the training set is generated using the synthetic tabular data generation (STDG) model. The evaluation framework is then applied to this synthetic dataset across 10 folds, simulating a cross-validation approach. Results from each fold are aggregated to compute the best, worst, and average values for various metrics and privacy risk measures. The entire process is executed in Python, utilizing the open-source MLflow platform for model tracking and result organization.
Results
The results section outlines the findings from the application of the proposed synthetic tabular data evaluation framework, as described in Section 2.2. The evaluation metrics and privacy risk measures for three datasets are summarized, emphasizing the influence of differential privacy (DP) on the synthetic tabular data generation (STDG) models. A detailed comparison of the performance of various STDG models is provided, with a focus on identifying the best-performing model.
Table 12 presents a comprehensive summary of the results across all evaluation categories and datasets, indicating the best and worst STDG models. The impact of DP is categorized as positive, negative, or neutral based on its effect on the majority of metrics or privacy risk measures, with significance determined at a threshold of $p < 0.05$. This classification allows for a nuanced understanding of how the incorporation of DP affects model performance and privacy outcomes. Further details are available in the Supplementary Material.
Discussion
The proposed synthetic tabular data evaluation framework emphasizes four key evaluation categories: fidelity, utility, privacy, and the tradeoff between fidelity and utility. These categories are critical for assessing how well synthetic data mimics real data (fidelity), its applicability in practical scenarios (utility), and its effectiveness in safeguarding sensitive information (privacy). The framework employs various metrics, including the Hellinger distance for fidelity, which measures the similarity between probability distributions of real and synthetic data, and the Pairwise Correlation Difference (PCD), which quantifies the preservation of correlations among variables. Additionally, the Depth vs. Depth plot (DD-Plot) and the Area Under the Receiver Operating Characteristic Curve (AUC-ROC) are utilized to evaluate multivariate distributional similarity and distinguishability between real and synthetic samples, respectively.
For utility assessment, the framework compares machine learning model performance using synthetic data against real data, employing metrics such as accuracy and mean absolute error. The fidelity-utility tradeoff is captured by the G metric, which balances fidelity and utility under varying privacy constraints. Privacy risk measures, including singling out, linkability, and membership inference risks, are also integrated to evaluate the effectiveness of synthetic data in protecting sensitive information. The framework’s application is validated using three health-related datasets, demonstrating its capability to analyze the impact of differential privacy on synthetic data quality and to identify the most suitable synthetic data generation models for medical applications.
Limitations
The limitations section of the paper highlights several areas for improvement in the proposed evaluation framework for synthetic tabular data generation (STDG) models, particularly in the context of fidelity, utility, and privacy. While these three categories are essential, the framework lacks additional evaluation metrics related to model efficiency, such as training time, resource consumption, generalization capability, and interpretability. Future research should aim to incorporate these metrics to provide a more comprehensive analysis of computational costs and to explore the interpretability of STDG models, which could enhance model selection for various applications.
Moreover, the methodology employed, which involved generating and evaluating 10 folds for validation, proved to be time-consuming. The authors suggest optimizing the computation of evaluation metrics and privacy risk measures, as well as employing multiple train-test splits to ensure representativeness and consistency. The reliance on default parameters for STDG models may not reflect their optimal performance; thus, hyperparameter tuning should be considered in future studies. Additionally, the impact of differential privacy (DP) on model performance was assessed using a single privacy budget, indicating a need for further exploration of varying privacy constraints and alternative privacy mechanisms. The authors also note specific models that require improvement, such as DP-TabDif, which performed poorly in fidelity and utility despite strong privacy results. Finally, the framework’s applicability could be extended to diverse datasets and domains, including federated environments, to assess its versatility and real-world utility in collaboration with clinical experts.
