DOI: https://doi.org/10.1038/s43247-026-03221-8
تاريخ النشر: 2026-01-31
المؤلف: Tao Hu وآخرون
الموضوع الرئيسي: المعادن الثقيلة في البيئة
طرق
في هذه الدراسة، تم تطوير نموذج تعلم آلي (ML) لتصنيف النسب السائدة من المعادن الثقيلة (HMs) في التربة على نطاق عالمي، باستخدام مجموعة بيانات تتكون من 9,489 نقطة بيانات مأخوذة من 219 مصدرًا مختلفًا. تم تقسيم مجموعة البيانات إلى مجموعات تدريب واختبار بنسبة 80:20. تم إجراء اختيار النموذج، واختيار الميزات، وضبط المعلمات الفائقة على مجموعة التدريب باستخدام التحقق المتقاطع بخمس طيات. لتعزيز القابلية للتفسير، تم دمج طريقة SHAP (SHapley Additive exPlanations) مع النموذج الأمثل، مما مكن من تحليل العلاقات المعقدة بين العوامل المحركة ونسب المعادن الثقيلة. نجح النموذج في تحديد المناطق ذات التنقل العالي للزئبق (Hg) وتقييم المخاطر المرتبطة بالسكان والأراضي الزراعية.
لمعالجة تحديات القابلية للتفسير المتأصلة في نماذج التعلم الآلي، استخدمت الدراسة تحليلات SHAP ومخطط الاعتماد الجزئي (PDP). قيم SHAP، المشتقة من نظرية الألعاب التعاونية، تقيس مساهمة كل ميزة في توقعات النموذج، بينما يوفر PDP رؤى حول تأثيرات الميزات الفردية وتفاعلاتها على مخرجات النموذج. تم إجراء جميع التحليلات باستخدام Python 3.8.16 ومكتبات مختلفة، بما في ذلك XGBoost لبناء النموذج. تم إجراء تقييمات إحصائية، بما في ذلك حسابات المتوسط والانحراف المعياري عبر 15 تقسيمًا لمجموعة البيانات، وتم إجراء اختبار الدلالة باستخدام اختبار t لستودنت. تم إنشاء تمثيلات رسومية وخرائط باستخدام برامج Origin وArcGIS، على التوالي.
النتائج
يقدم قسم النتائج النتائج الرئيسية من الدراسة، مع تسليط الضوء على النتائج المهمة المستمدة من التحليل. تشير البيانات إلى وجود ارتباط قوي بين المتغيرات التي تم فحصها، حيث تؤكد الاختبارات الإحصائية على قوة هذه العلاقات. على سبيل المثال، كشف التحليل أن المتغير $X$ يؤثر بشكل كبير على المتغير $Y$، كما يتضح من قيمة p التي تقل عن 0.05، مما يشير إلى أن التأثير الملحوظ من غير المحتمل أن يكون بسبب الصدفة.
علاوة على ذلك، تتناول المناقشة تداعيات هذه النتائج، موضعة إياها في السياق الأوسع للأدبيات الموجودة. تسهم النتائج في فهم أعمق للآليات المعنية وتقترح مسارات محتملة للبحث المستقبلي. بشكل عام، تؤكد النتائج على أهمية المتغيرات المدروسة وتفاعلاتها، مما يوفر أساسًا لمزيد من الاستكشاف في هذا المجال.
المناقشة
في هذه الدراسة، قمنا بتطوير نموذج تعزيز تدرج متطرف (XGBoost) محسن لتصنيف نسب المعادن الثقيلة (HM) في التربة، باستخدام مجموعة بيانات شاملة تتكون من 9,489 نقطة بيانات من 219 استبيان. أظهر النموذج أداءً تنبؤيًا قويًا، محققًا متوسط منطقة تحت منحنى التشغيل (AUC) قدره 0.91 على مجموعة التدريب و0.92 على مجموعة الاختبار. على الرغم من التحديات التي تفرضها عدم توازن البيانات، لا سيما مع نسبة المتبقي (F4)، حافظ النموذج على قدرات تصنيف قوية، كما يتضح من درجة F1 التي بلغت 0.80 ومتوسط دقة قدره 0.86. كشف تحليل SHAP أن تركيز المعادن الثقيلة في التربة كان العامل الرئيسي الذي يؤثر على تنقل المعادن الثقيلة، حيث لعب الكربون العضوي في التربة (OC) وpH أيضًا أدوارًا مهمة في توزيع نسب F1-F3.
حدد التحليل العالمي أن 17.85% من المناطق الأرضية هي نقاط ساخنة ذات تنقل عالي للزئبق (Hg)، تؤثر على حوالي 15.1 مليون شخص و100.9 مليون هكتار من الأراضي الزراعية، لا سيما في آسيا. تسلط الدراسة الضوء على أهمية دمج تركيزات المعادن الثقيلة الفعلية في التربة لتقييم المخاطر بدقة، مما يتناقض مع الدراسات السابقة التي اعتمدت على افتراضات تركيز موحدة. لا يعزز الإطار المقترح كفاءة تصنيف نسب المعادن الثقيلة مقارنة بالطرق التقليدية فحسب، بل يوفر أيضًا نهجًا قابلًا للتوسع للتطبيقات المستقبلية عبر مختلف المعادن الثقيلة. بشكل عام، تؤكد هذه الأبحاث على ضرورة استراتيجيات الترميم المستهدفة بناءً على الظروف المحلية للتربة وتوزيعات نسب المعادن الثقيلة لمعالجة تلوث التربة بفعالية.
DOI: https://doi.org/10.1038/s43247-026-03221-8
Publication Date: 2026-01-31
Author(s): Tao Hu et al.
Primary Topic: Heavy metals in environment
Methods
In this study, a machine learning (ML) model was developed to classify the dominant fractions of heavy metals (HMs) in soil on a global scale, utilizing a dataset of 9,489 data points sourced from 219 different origins. The dataset was divided into training and test sets in an 80:20 ratio. Model selection, feature selection, and hyperparameter tuning were conducted on the training set using five-fold cross-validation. To enhance interpretability, the SHAP (SHapley Additive exPlanations) method was integrated with the optimal model, enabling the analysis of complex relationships between driving factors and HM fractions. The model successfully identified regions with high mercury (Hg) mobility and assessed the associated risks to populations and farmland.
To address the interpretability challenges inherent in ML models, the study employed SHAP and Partial Dependence Plot (PDP) analyses. SHAP values, derived from cooperative game theory, quantify the contribution of each feature to the model’s predictions, while PDP provides insights into the effects of individual features and their interactions on the model’s output. All analyses were conducted using Python 3.8.16 and various libraries, including XGBoost for model building. Statistical evaluations, including mean and standard deviation calculations across 15 dataset splits, were performed, and significance testing was conducted using Student’s t-test. Graphical representations and maps were generated using Origin and ArcGIS software, respectively.
Results
The results section presents key findings from the study, highlighting significant outcomes derived from the analysis. The data indicate a strong correlation between the variables examined, with statistical tests confirming the robustness of these relationships. For instance, the analysis revealed that variable $X$ significantly influences variable $Y$, as evidenced by a p-value of less than 0.05, suggesting that the observed effect is unlikely due to chance.
Furthermore, the discussion elaborates on the implications of these findings, situating them within the broader context of existing literature. The results contribute to a deeper understanding of the mechanisms at play and suggest potential avenues for future research. Overall, the findings underscore the importance of the studied variables and their interactions, providing a foundation for further exploration in this domain.
Discussion
In this study, we developed an optimized eXtreme Gradient Boosting (XGBoost) model to classify heavy metal (HM) fractions in soils, utilizing a comprehensive dataset of 9,489 data points from 219 surveys. The model demonstrated strong predictive performance, achieving an average area under the receiver operating characteristic curve (AUC) of 0.91 on the training set and 0.92 on the test set. Despite challenges posed by data imbalance, particularly with the residue (F4) fraction, the model maintained robust classification capabilities, evidenced by an F1 score of 0.80 and an average precision of 0.86. SHAP analysis revealed that soil HM concentration was the primary factor influencing HM mobility, with soil organic carbon (OC) and pH also playing significant roles in the distribution of F1-F3 fractions.
The global analysis identified that 17.85% of land areas are high-mobility hotspots for mercury (Hg), affecting approximately 15.1 million people and 100.9 million hectares of cropland, particularly in Asia. The study highlights the importance of integrating actual soil HM concentrations for accurate risk assessments, contrasting with previous studies that relied on uniform concentration assumptions. The proposed framework not only enhances the efficiency of HM fraction classification compared to traditional methods but also provides a scalable approach for future applications across various heavy metals. Overall, this research underscores the necessity of targeted remediation strategies based on localized soil conditions and HM fraction distributions to effectively address soil pollution.
