DOI: https://doi.org/10.1038/s41598-026-40062-z
PMID: https://pubmed.ncbi.nlm.nih.gov/41896256
تاريخ النشر: 2026-03-27
المؤلف: Mohamed Yasser AboElsaad وآخرون
الموضوع الرئيسي: تطبيقات الشبكات العصبية المتقدمة
نظرة عامة
تقدم البحث KM-DBSCAN، وهو خوارزمية تجميع جديدة مصممة لتعزيز تقليل البيانات في تعلم الآلة مع الالتزام بمبادئ الذكاء الاصطناعي الأخضر. من خلال دمج الكفاءة الهندسية لـ K-Means مع الوعي بالكثافة لـ DBSCAN، تحقق KM-DBSCAN تقليلًا كبيرًا في البيانات – يصل إلى 90% – وتسريع التدريب يتراوح من 3.6× إلى 6900× عبر نماذج مختلفة (SVM، MLP، CNN) ومجموعات بيانات (موز، USPS، Adult9a، Collision، Dry Bean، Melanoma). ومن الجدير بالذكر أنها تحافظ على دقة تنافسية، كما يتضح من دقة تصنيف الميلانوما التي بلغت 90.39% باستخدام 28.7% فقط من بيانات التدريب، مما أدى إلى فقدان دقة ضئيل قدره 0.0061% وتقليل كبير في انبعاثات الكربون.
تعمل الطريقة من خلال تطبيق K-Means أولاً لتوليد المراكز، مما يقلل من التعقيد الحسابي لـ DBSCAN من \(O(n^2)\) إلى \(O(k^2)\)، مما يعزز من قوة المعلمات ويحسن من اكتشاف الحدود في التجمعات المتداخلة. تظهر النتائج التجريبية أن KM-DBSCAN تتفوق باستمرار على الطرق الحالية من حيث نسبة التخفيض والمعايير البيئية، مما يجعلها مناسبة بشكل خاص للبيئات ذات الموارد المحدودة مثل أنظمة الذكاء الاصطناعي على الحافة والتشخيصات الطبية. بشكل عام، لا تعالج KM-DBSCAN فقط قيود طرق التجميع التقليدية ولكنها توفر أيضًا إطارًا مستدامًا لتعلم الآلة الفعال، مما يمهد الطريق لاعتماد أنظمة الذكاء الاصطناعي المدروسة بيئيًا.
مقدمة
تناقش المقدمة أهمية التجميع غير المراقب في تعلم الآلة، مع التركيز على دوره في تجميع البيانات غير المعلّمة بناءً على التشابه. تبرز خوارزميتين رئيسيتين: K-Means و DBSCAN. تهدف K-Means إلى تقليل تباين داخل التجمع من خلال تقسيم البيانات إلى $k$ تجمعات، لكنها محدودة بفرضياتها حول التجمعات المحدبة والكرية وحساسيتها للافتتاح. بالمقابل، تحدد DBSCAN التجمعات بناءً على المناطق الكثيفة، مما يسمح بأشكال عشوائية ومرونة تجاه الضوضاء، على الرغم من أنها تواجه تحديات مع الكثافات المتغيرة والحدود المتداخلة.
تستكشف هذه الفقرة أيضًا تقنيات تقليل البيانات المختلفة التي تعزز من كفاءة نماذج تعلم الآلة، لا سيما في تدريب آلات الدعم (SVM). تشمل الطرق البارزة نهج التخفيض ذو المرحلتين لشين وآخرين، الذي يقوم بتصفية النقاط غير الأساسية بناءً على القرب من مراكز التجمع، وخوارزمية التجميع القائمة على القشرة المحدبة (CBCH) لبيرزهاندي ويون، التي تستخدم K-Means للتقسيم وQuickhull لتقليل التجمعات الفردية. تركز تقنيات أخرى، مثل خوارزمية استخراج النقاط الحدودية وطريقة استخراج القشرة، على الاحتفاظ بالنقاط الحدودية المعلوماتية مع القضاء على البيانات الزائدة. تختتم المقدمة بالإشارة إلى التركيز المتزايد على الكفاءة الحسابية والاستدامة في الذكاء الاصطناعي، مما يضع تركيز الدراسة الحالية على تقليل البيانات على مستوى الحالة كنهج تكميلي لتحسينات على مستوى النموذج الحالية.
طرق
في هذا القسم، يوضح المؤلفون إعداد التجارب والنتائج التي تم الحصول عليها من تطبيق طريقتهم المقترحة KM-DBSCAN عبر مجموعات بيانات مرجعية مختلفة. كانت التجارب تهدف إلى تقييم فعالية الطريقة في التجميع وتقليل البيانات مع الحفاظ على دقة تصنيف عالية. تم إجراء التجارب باستخدام Python 3.10.4 على محطة عمل محلية مزودة بمعالج Intel Core i7-10750H، و بطاقة رسومات NVIDIA GeForce GTX 1660 Ti، و 16 جيجابايت من الذاكرة العشوائية، تعمل على نظام Windows 10. استند التنفيذ إلى مكتبات scikit-learn و CodeCarbon لمعالجة البيانات وتتبع انبعاثات الكربون، على التوالي.
لمراقبة انبعاثات الكربون، تم استخدام متتبع CodeCarbon (v2.3.5) في وضع الآلة لتسجيل استهلاك الطاقة الفعلي لوحدة المعالجة المركزية/وحدة معالجة الرسومات وتقدير الانبعاثات بالكيلوغرام من مكافئ CO₂. تم استخدام عامل كثافة الكربون الافتراضي لمنطقة مصر/شمال إفريقيا لضمان تقدير دقيق للانبعاثات. تم تفعيل المتتبع فقط خلال مرحلة تدريب المصنف، مما يعزل استهلاك الطاقة المرتبط بتدريب النموذج عن الأنشطة غير النشطة أو ما بعد المعالجة. تضمن هذه الطريقة المنهجية تقديرات شفافة وقابلة للتكرار لاستخدام الطاقة والانبعاثات المرتبطة بالتجارب.
نقاش
يعالج إطار العمل المقترح KM-DBSCAN التحديات المتعلقة بتقليل البيانات في تعلم الآلة من خلال دمج نقاط القوة في طرق التجميع K-Means و DBSCAN. هذه الطريقة الهجينة تقلل بشكل فعال من التكاليف الحسابية مع الحفاظ على دقة النموذج وتقليل الأثر البيئي، مما يتماشى مع مبادئ الذكاء الاصطناعي الأخضر. تستخدم الطريقة عملية من مرحلتين: أولاً، يتم استخدام K-Means لضغط مجموعة البيانات إلى $k$ مراكز، مما يبسط مشكلة التجميع لـ DBSCAN، ويقلل من تعقيد الوقت من $O(n^2)$ إلى $O(k^2)$. يسمح هذا الخطوة المسبقة لـ DBSCAN بالعمل على مجموعة أكثر قابلية للإدارة من المراكز، مما يعزز من كفاءتها ومرونتها، لا سيما في السيناريوهات التي تحتوي على توزيعات صفية متداخلة.
تظهر خوارزمية KM-DBSCAN تحسينات كبيرة في جودة التجميع وكفاءة الحساب عبر مجموعات بيانات مختلفة. تشير النتائج التجريبية إلى أن الطريقة تحقق درجات عالية من مؤشر ران المعدل (ARI) بينما تقلل بشكل كبير من وقت التشغيل. على سبيل المثال، في مجموعة بيانات Two Moons، حافظت KM-DBSCAN على ARI قدره 0.976 باستخدام 100 نقطة تمثيلية فقط، مما أدى إلى تسريع قدره 23.9× مقارنة بـ DBSCAN التقليدية. علاوة على ذلك، تقلل الإطار بشكل فعال من حجم البيانات مع الحفاظ على الخصائص الأساسية لتوزيع الصفوف، مما يجعلها مناسبة بشكل خاص للتطبيقات واسعة النطاق. إن قدرة الطريقة على تحقيق التوازن بين الدقة والكفاءة، إلى جانب مرونتها تجاه ضبط المعلمات، تجعلها حلاً واعدًا لتقليل البيانات الذكي في مهام تعلم الآلة.
القيود
يعترف قسم القيود بعدة قيود لخوارزمية KM-DBSCAN على الرغم من كفاءتها وتحسينات أدائها المثبتة. أولاً، بينما تكون KM-DBSCAN أقل حساسية لاختيارات المعلمات مقارنة بـ DBSCAN التقليدية، لا تزال المعلمات مثل عدد التجمعات \( k \) في K-Means ومعلمات \( \epsilon \)/MinPts في DBSCAN تتطلب ضبطًا محددًا لمجموعات البيانات، مما يبرز الحاجة إلى اختيار المعلمات التلقائي في الأبحاث المستقبلية. ثانيًا، تعتمد فعالية تقليل الحالة على توزيع البيانات؛ قد تؤثر الحدود الصفية غير المنتظمة أو المتداخلة سلبًا على جودة التخفيض.
بالإضافة إلى ذلك، يثير الاعتماد على ضغط المراكز واختيار الحدود المستند إلى الكثافة مخاوف بشأن إمكانية تمثيل العينات من الفئات الأقل أو النادرة بشكل غير كافٍ في مجموعات البيانات غير المتوازنة. تهدف الأعمال المستقبلية إلى معالجة هذه القيود من خلال تعزيز قوة ومرونة وقدرات تعميم إطار KM-DBSCAN. يقترح المؤلفون أيضًا استكشاف قابليته للتطبيق عبر مجالات مختلفة تتجاوز التصوير الطبي، مثل دمج التخفيض القائم على الكثافة مع نماذج الانتباه البصري لتحسين مهام معالجة الصور، وتكييف KM-DBSCAN لتحليلات البيانات واسعة النطاق في الأمن السيبراني وأنظمة إنترنت الأشياء، مما يعزز الاستدامة في التعامل مع مجموعات البيانات المعقدة.
DOI: https://doi.org/10.1038/s41598-026-40062-z
PMID: https://pubmed.ncbi.nlm.nih.gov/41896256
Publication Date: 2026-03-27
Author(s): Mohamed Yasser AboElsaad et al.
Primary Topic: Advanced Neural Network Applications
Overview
The research introduces KM-DBSCAN, a novel clustering algorithm designed to enhance data reduction in machine learning while adhering to Green AI principles. By integrating the geometric efficiency of K-Means with the density-awareness of DBSCAN, KM-DBSCAN achieves significant data reduction—up to 90%—and training speedups ranging from 3.6× to 6900× across various models (SVM, MLP, CNN) and datasets (Banana, USPS, Adult9a, Collision, Dry Bean, Melanoma). Notably, it maintains competitive accuracy, exemplified by a melanoma classification accuracy of 90.39% using only 28.7% of the training data, resulting in a minimal accuracy loss of 0.0061% and a substantial reduction in carbon emissions.
The method operates by first applying K-Means to generate centroids, which reduces the computational complexity of DBSCAN from \(O(n^2)\) to \(O(k^2)\), thereby enhancing parameter robustness and improving border detection in overlapping clusters. Experimental results demonstrate that KM-DBSCAN consistently outperforms existing methods in terms of reduction ratio and environmental metrics, making it particularly suitable for resource-constrained environments such as edge-AI systems and medical diagnostics. Overall, KM-DBSCAN not only addresses the limitations of traditional clustering methods but also provides a sustainable framework for efficient machine learning, paving the way for the adoption of environmentally conscious AI systems.
Introduction
The introduction discusses the significance of unsupervised clustering in machine learning, emphasizing its role in grouping unlabeled data based on similarity. It highlights two key algorithms: K-Means and DBSCAN. K-Means aims to minimize intra-cluster variance by partitioning data into $k$ clusters, but it is limited by its assumptions of convex, spherical clusters and sensitivity to initialization. In contrast, DBSCAN identifies clusters based on dense regions, allowing for arbitrary shapes and robustness to noise, though it faces challenges with varying densities and overlapping boundaries.
The section further explores various data reduction techniques that enhance the efficiency of machine learning models, particularly in support vector machine (SVM) training. Notable methods include Shen et al.’s two-stage reduction approach, which filters out non-essential points based on proximity to cluster centroids, and Birzhandi and Youn’s Clustering-Based Convex Hull (CBCH) algorithm, which utilizes K-Means for partitioning and Quickhull for reducing singular clusters. Other techniques, such as the Border Point extraction algorithm and Shell Extraction method, focus on retaining informative boundary points while eliminating redundant data. The introduction concludes by noting the increasing emphasis on computational efficiency and sustainability in AI, positioning the current study’s focus on instance-level data reduction as a complementary approach to existing model-level optimizations.
Methods
In this section, the authors detail the experimental setup and results obtained from the application of their proposed KM-DBSCAN method across various benchmark datasets. The experiments aimed to assess the method’s effectiveness in clustering and data reduction while preserving high classification accuracy. The experiments were conducted using Python 3.10.4 on a local workstation with an Intel Core i7-10750H CPU, an NVIDIA GeForce GTX 1660 Ti GPU, and 16 GB of RAM, operating on Windows 10. The implementation leveraged the scikit-learn and CodeCarbon libraries for data processing and carbon emission tracking, respectively.
To monitor carbon emissions, the CodeCarbon tracker (v2.3.5) was employed in machine-mode to record real-time CPU/GPU power consumption and estimate emissions in kg CO₂ equivalent. The default carbon-intensity factor for the Egypt/North Africa region was utilized to ensure accurate emissions estimation. The tracker was activated solely during the classifier training phase, thereby isolating energy consumption related to model training from idle or post-processing activities. This methodological approach guarantees transparent and reproducible estimates of energy usage and emissions associated with the experiments.
Discussion
The proposed KM-DBSCAN framework addresses the challenges of data reduction in machine learning by combining the strengths of K-Means and DBSCAN clustering methods. This hybrid approach effectively reduces computational costs while maintaining model accuracy and minimizing environmental impact, aligning with the principles of Green AI. The method employs a two-stage process: first, K-Means is used to compress the dataset into $k$ centroids, which simplifies the clustering problem for DBSCAN, reducing its time complexity from $O(n^2)$ to $O(k^2)$. This preprocessing step allows DBSCAN to operate on a more manageable set of centroids, enhancing its efficiency and robustness, particularly in scenarios with overlapping class distributions.
The KM-DBSCAN algorithm demonstrates significant improvements in clustering quality and computational efficiency across various datasets. Empirical results indicate that the method achieves high Adjusted Rand Index (ARI) scores while drastically reducing runtime. For instance, in the Two Moons dataset, KM-DBSCAN maintained an ARI of 0.976 using only 100 representative points, resulting in a 23.9× speedup compared to traditional DBSCAN. Furthermore, the framework effectively reduces data size while preserving critical class distribution characteristics, making it particularly suitable for large-scale applications. The method’s ability to balance accuracy and efficiency, alongside its robustness to parameter tuning, positions it as a promising solution for intelligent data reduction in machine learning tasks.
Limitations
The section on limitations acknowledges several constraints of the KM-DBSCAN algorithm despite its demonstrated efficiency and performance improvements. Firstly, while KM-DBSCAN is less sensitive to hyperparameter choices than traditional DBSCAN, parameters such as the number of clusters \( k \) in K-Means and the \( \epsilon \)/MinPts parameters in DBSCAN still necessitate dataset-specific tuning, highlighting the need for automated hyperparameter selection in future research. Secondly, the effectiveness of instance reduction is contingent upon the data distribution; irregular or overlapping class boundaries may adversely affect the quality of reduction.
Additionally, the reliance on centroid-based compression and density-derived border selection raises concerns about the potential under-representation of minority or rare-class samples in imbalanced datasets. Future work aims to address these limitations by enhancing the robustness, fairness, and generalization capabilities of the KM-DBSCAN framework. The authors also propose to explore its applicability across various domains beyond medical imaging, such as integrating density-based reduction with visual attention models to optimize image processing tasks, and adapting KM-DBSCAN for large-scale data analytics in cybersecurity and IoT systems, thereby promoting sustainability in handling complex datasets.
