استخراج الميزات في مجموعات بيانات أمراض النباتات الحسية باستخدام وظائف العضوية المعاد تشكيلها المستقلة عن متغيرات الفئة
Feature extraction in sensor plant disease datasets using reformed membership functions independent of class variables

شارك:
المجلة: Scientific Reports، المجلد: 16، العدد: 1
DOI: https://doi.org/10.1038/s41598-025-33569-4
PMID: https://pubmed.ncbi.nlm.nih.gov/41617750
تاريخ النشر: 2026-01-30
المؤلف: Ayushi Gupta وآخرون
الموضوع الرئيسي: الزراعة الذكية والذكاء الاصطناعي

نظرة عامة

تقدم البحث تقنية استخراج الميزات المعتمدة على دالة العضوية (MFFE) التي تهدف إلى تحسين مجموعات بيانات النباتات المعتمدة على المستشعرات ذات الحجم الصغير، وخاصة لاكتشاف أمراض الطماطم. من خلال استخدام مجموعتين من البيانات المعتمدة على المستشعرات في الوقت الحقيقي، TomEBD و TPMD، يتناول الدراسة عدم توازن البيانات من خلال تقنية KMeans-SMOTE. تستخدم استخراج الميزات دوال عضوية مثلثية وغوسية مُعدلة، مما يضمن أن جميع المعلمات مشتقة فقط من بيانات التدريب لتجنب تسرب المعلومات. يتم تصنيف مجموعات البيانات المحسنة باستخدام نموذجين محسّنين: آلة التعلم المتطرفة المعتمدة على النواة المحسّنة (OKELM) وشبكة الأعصاب المعتمدة على دالة الأساس الشعاعي المحسّنة (ORBFNN)، وكلاهما محسّن عبر إطار عمل Optuna. أظهر نموذج TMF-ORBFNN دقة متفوقة عبر كل من مجموعات بيانات النباتات والأمراض ومجموعات البيانات المرجعية، تم التحقق منها من خلال اختبارات إحصائية تشير إلى اختلافات أداء كبيرة عن النماذج الأخرى.

تؤكد الدراسة على أهمية استخراج الميزات قبل التصنيف في مجموعات البيانات المعتمدة على المستشعرات الصغيرة وتبرز قوة وقدرات تعميم نموذج TMF-ORBFNN. بينما تقلل الطريقة المقترحة من التحيز الإحصائي واعتماد المتغيرات الفئوية، تواجه قيودًا مثل أوقات المعالجة الأبطأ والضوضاء المحتملة في بيانات المستشعر. يُقترح أن تركز الأعمال المستقبلية على تطوير مستشعرات منخفضة التكلفة، وتعزيز الكفاءة الحسابية، وتوسيع الإطار لأنظمة التنبؤ بأمراض المحاصيل المتعددة. بالإضافة إلى ذلك، يُقترح تطوير تطبيق لمراقبة الظروف الزراعية في الوقت الحقيقي لمساعدة المزارعين في إدارة الأمراض في الوقت المناسب.

طرق

تم هيكلة قسم المنهجية في الدراسة إلى سبع فقرات فرعية، كل منها تتناول مكونات حاسمة من البحث. تتناول الفقرة الفرعية “إجراء جمع البيانات” تفاصيل الحصول على مجموعتين من بيانات أمراض النباتات جنبًا إلى جنب مع ثماني مجموعات بيانات مرجعية، مما يضع الأساس للتحليل. توضح الفقرة الفرعية “إعادة أخذ العينات” تقنيات إعادة أخذ العينات المستخدمة لتعزيز قوة البيانات.

في الفقرة الفرعية “استخراج الميزات المعتمد على دالة العضوية (MFFE)”، يقدم المؤلفون نهجهم الجديد لاستخراج الميزات، والذي يعد محوريًا للتحليل اللاحق. تصف الفقرة الفرعية “مصنفات التعلم الآلي” المصنفين المعتمدين على التعلم الآلي المستخدمين في الدراسة، بينما تناقش “تحسين المعلمات الفائقة” عمليات الضبط المطبقة لتحسين أداء النموذج. تلخص الفقرة الفرعية “التجارب الأولية” التجارب الأولية التي شكلت الإطار التجريبي، لتصل إلى الفقرة الفرعية “إطار البحث”، التي تقدم الهيكل العام الذي يوجه البحث.

نتائج

تم هيكلة قسم “النتائج” في الدراسة إلى سبع فقرات فرعية، كل منها تتناول جوانب مختلفة من أداء النموذج والتحليل. تتناول الفقرتان الفرعيتان بعنوان “الأداء على مجموعات بيانات النباتات” و “الأداء على مجموعات البيانات المرجعية” فعالية ستة نماذج تم تقييمها على مجموعتين متميزتين من بيانات النباتات وثماني مجموعات بيانات مرجعية، على التوالي. تقدم الفقرة الفرعية “اختبار فريدمان واختبار بونفيروني-دون بعد ذلك” تحليلًا إحصائيًا للتحقق من اختلافات الأداء بين النماذج، مما يضمن مقارنات قوية.

بالإضافة إلى ذلك، تناقش الفقرة الفرعية “أهمية المعلمات الفائقة” دور المعلمات الفائقة في تحسين دقة النموذج، بينما تفحص الفقرة الفرعية “تحليل تعقيد الوقت” الكفاءة الحسابية للنماذج. تضع الفقرة الفرعية “المقارنة مع الدراسات الحالية” التقنية المقترحة مقابل الأساليب الحديثة الحالية، مبرزة مزاياها النسبية. تختتم القسم بـ “نقاش” يلخص النتائج وآثارها.

نقاش

تحدد قسم النقاش في ورقة البحث فجوة كبيرة في الأدبيات الحالية حول اكتشاف أمراض النباتات، خاصة فيما يتعلق بالتركيز المحدود على استخراج الميزات من البيانات المعتمدة على المستشعرات. بينما استخدمت العديد من الدراسات طرقًا تعتمد على الصور، يبرز المؤلفون أن بيانات المستشعر غالبًا ما تفتقر إلى ميزات كافية بسبب التكاليف العالية المرتبطة بواجهات المستشعر. قدم العمل السابق من Bhatia و Chug طريقة لاستخراج الميزات تعتمد على الفئة باستخدام دالة عضوية مثلثية ضبابية معدلة، لكن هذه الطريقة لها قيود ملحوظة، بما في ذلك اعتمادها على المتغيرات الفئوية أثناء التدريب، مما يعيق التطبيق في الوقت الحقيقي ويؤدي إلى نتائج متحيزة.

لمعالجة هذه القضايا، تقترح الدراسة الحالية تقنية جديدة لاستخراج الميزات المعتمدة على دالة العضوية (MFFE) التي تقضي على اعتماد المتغيرات الفئوية وتقلل من التحيز الإحصائي من خلال حساب معلمات دالة العضوية فقط على بيانات التدريب. تسمح هذه الطريقة بتحويل الميزات في سيناريوهات الوقت الحقيقي حيث تكون معلومات الفئة غير متاحة. يتحقق المؤلفون من نهجهم باستخدام مجموعتين من بيانات أمراض الطماطم وثماني مجموعات بيانات مرجعية، باستخدام مصنفات التعلم الآلي المحسّنة من خلال إطار عمل Optuna. تشمل مساهمات الدراسة تقديم طريقة لاستخراج الميزات غير المعتمدة على الفئة، ومنع تسرب المعلومات، وتطوير نماذج محسّنة لتنبؤ الأمراض، بهدف تعزيز الممارسات الزراعية من خلال تمكين تدابير التحكم في الأمراض في الوقت المناسب.

Journal: Scientific Reports, Volume: 16, Issue: 1
DOI: https://doi.org/10.1038/s41598-025-33569-4
PMID: https://pubmed.ncbi.nlm.nih.gov/41617750
Publication Date: 2026-01-30
Author(s): Ayushi Gupta et al.
Primary Topic: Smart Agriculture and AI

Overview

The research presents a Membership Function-based Feature Extraction (MFFE) technique aimed at enhancing small-sized sensor-based plant datasets, particularly for tomato disease detection. By utilizing two real-time sensor-based datasets, TomEBD and TPMD, the study addresses dataset imbalance through the KMeans-SMOTE technique. Feature extraction employs reformed triangular and Gaussian membership functions, ensuring that all parameters are derived solely from training data to avoid information leakage. The enhanced datasets are classified using two optimized models: the Optimized Kernel Extreme Learning Machine (OKELM) and the Optimized Radial Basis Function Neural Network (ORBFNN), both optimized via the Optuna framework. The TMF-ORBFNN model demonstrated superior accuracy across both plant-disease and benchmark datasets, validated through statistical tests indicating significant performance differences from other models.

The study emphasizes the importance of feature extraction prior to classification in small sensor-based datasets and highlights the robustness and generalization capabilities of the TMF-ORBFNN model. While the proposed method effectively reduces statistical bias and class-variable dependency, it faces limitations such as slower processing times and potential noise in sensor data. Future work is suggested to focus on developing low-cost sensors, enhancing computational efficiency, and expanding the framework for multi-crop disease prediction systems. Additionally, the development of a real-time monitoring application for agricultural conditions is proposed to aid farmers in timely disease management.

Methods

The methodology section of the study is structured into seven subsections, each addressing critical components of the research. The “Data collection procedure” subsection details the acquisition of two plant disease datasets alongside eight benchmarking datasets, establishing the foundation for the analysis. The “Data resampling” subsection outlines the resampling techniques employed to enhance the robustness of the data.

In the “Membership function based feature extraction (MFFE)” subsection, the authors introduce their novel feature extraction approach, which is pivotal for the subsequent analysis. The “ML classifiers” subsection describes the two machine learning classifiers utilized in the study, while “Hyperparameter optimization” discusses the tuning processes applied to optimize model performance. The “Preliminary experiments” subsection summarizes initial experiments that shaped the experimental framework, culminating in the “Research framework” subsection, which presents the overarching structure guiding the research.

Results

The “Results” section of the study is structured into seven subsections, each addressing different aspects of the model performance and analysis. The subsections titled “Performance on plant datasets” and “Performance on benchmarking datasets” detail the efficacy of six models evaluated on two distinct plant datasets and eight benchmarking datasets, respectively. The “Friedman and post-hoc Bonferroni-Dunn test” subsection presents a statistical analysis to verify performance differences among the models, ensuring robust comparisons.

Additionally, the “Hyperparameter importance” subsection discusses the role of hyperparameters in optimizing model accuracy, while the “Time complexity analysis” subsection examines the computational efficiency of the models. The “Comparison with the existing studies” subsection positions the proposed technique against current state-of-the-art methods, highlighting its relative advantages. The section concludes with a “Discussion” that synthesizes the findings and implications of the results.

Discussion

The discussion section of the research paper identifies a significant gap in the existing literature on plant disease detection, particularly regarding the limited focus on feature extraction from sensor-based data. While many studies have utilized image-based methods, the authors highlight that sensor data often lacks sufficient features due to the high costs associated with sensor interfaces. The prior work by Bhatia and Chug introduced a class-dependent feature extraction method using a modified triangular fuzzy membership function, but this approach has notable limitations, including its reliance on class variables during training, which hinders real-time application and leads to biased results.

To address these issues, the current study proposes a novel Membership Function-based Feature Extraction (MFFE) technique that eliminates class variable dependency and reduces statistical bias by computing membership function parameters solely on the training data. This method allows for the transformation of features in real-time scenarios where class information is unavailable. The authors validate their approach using two tomato disease datasets and eight benchmark datasets, employing machine learning classifiers optimized through the Optuna framework. The study’s contributions include the introduction of a class-independent feature extraction method, prevention of information leakage, and the development of optimized models for disease prediction, ultimately aiming to enhance agricultural practices by enabling timely disease control measures.

شارك: