DOI: https://doi.org/10.1017/qpb.2025.10018
PMID: https://pubmed.ncbi.nlm.nih.gov/40735612
تاريخ النشر: 2025-01-01
المؤلف: Kaiyue Hong وآخرون
الموضوع الرئيسي: الزراعة الذكية والذكاء الاصطناعي
نظرة عامة
لقد أثرت التطورات الحديثة في علم البيانات والذكاء الاصطناعي بشكل كبير على علوم النبات، وخاصة من خلال تطبيق تقنيات التعرف على الصور والتعلم العميق. تسلط هذه المراجعة الضوء على التأثيرات التحويلية لهذه الابتكارات على تحديد الأنواع، واكتشاف الأمراض، وتحليل الإشارات الخلوية، ومراقبة النمو. وتؤكد على الدور الحاسم لاكتساب البيانات ومعالجتها مسبقًا، موضحة تقنيات مثل التصوير عالي الدقة وتصوير الطائرات بدون طيار، بالإضافة إلى طرق تحسين الصور مثل القص والتكبير.
علاوة على ذلك، تتناول المراجعة تقنيات استخراج الميزات، بما في ذلك هيستوجرامات الألوان وتحليل القوام، والتي تعتبر حيوية لتحديد النباتات بشكل فعال وتقييم صحتها. يستكشف المؤلفون أيضًا الاتجاهات الناشئة، والتحديات، والاتجاهات المستقبلية في هذا المجال، موفرين رؤى حول كيفية تقدم هذه الأدوات الحاسوبية كل من أبحاث علوم النبات والتطبيقات العملية.
مقدمة
تؤكد مقدمة الورقة على الدور الحاسم لمجموعات بيانات صور النباتات على نطاق واسع في تقدم علوم النبات، بينما تتناول أيضًا التحديات المرتبطة بمعالجتها بكفاءة. تسلط الضوء على الإمكانات التحويلية للذكاء الاصطناعي (AI) والتعلم العميق (DL) في أتمتة تحليل صور النباتات من خلال تقنيات التعلم الآلي (ML). تمكن هذه التقنيات الحواسيب من التعلم من البيانات، وتحديد الأنماط، وإجراء التنبؤات، مما يعزز بشكل كبير المهام مثل مراقبة نمو النباتات واكتشاف الأمراض. تشير الورقة إلى الصعوبات التي تطرحها تنوع النباتات والخلفيات المعقدة في معالجة الصور، مما يتطلب طرق معالجة مسبقة مخصصة، واستخراج الميزات، واستراتيجيات زيادة البيانات لتحسين أداء النموذج.
علاوة على ذلك، تناقش المقدمة التقدم في التعرف على الصور الذي تسهله الشبكات العصبية التلافيفية (CNNs)، والتي أظهرت دقة عالية في التعرف على أنواع النباتات عبر مجموعات بيانات مختلفة، متفوقة على الطرق التقليدية. كما تقدم إمكانات نماذج اللغة الكبيرة (LLMs) مثل ChatGPT وAgronomic Nucleotide Transformer (AgroNT) في علم الوراثة النباتية، موضحة قدرتها على تحليل البيانات الجينومية واكتشاف رؤى جديدة حول التعبير الجيني واستجابات الإجهاد. تهدف المراجعة إلى تقييم التقنيات الرئيسية في معالجة صور النباتات، مقارنة المنهجيات ونماذج ML لتوفير إطار شامل للباحثين في هذا المجال.
طرق
تحدد قسم الطرق المكونات الحاسمة لتدريب وتقييم نموذج التعلم الآلي (ML) بشكل فعال، مع التأكيد على تقسيم البيانات، ودوال الخسارة، واستراتيجيات التحسين. يُوصى بتقسيم البيانات القياسي بنسبة 70:15:15 للتدريب، والتحقق، والاختبار لضمان تعميم النموذج بشكل قوي. تُظهر دراسات بارزة، مثل تلك التي أجراها موهانتي وآخرون وفيرنتينوس، فعالية التعلم العميق في تحديد أمراض النباتات وتصنيف الأنواع، محققة دقة تتجاوز 99%. يناقش القسم أيضًا مختلف دوال الخسارة المخصصة لمهام الانحدار والتصنيف، مع تسليط الضوء على أهمية دوال الخسارة المخصصة، كما يتضح من عمل جيلسبي وآخرين في مراقبة الأنواع النادرة من النباتات.
تلعب خوارزميات التحسين دورًا محوريًا في تعزيز أداء النموذج، مع تقنيات شائعة تشمل الانحدار العشوائي (SGD)، وتقدير الزخم التكيفي (ADAM)، وRMSprop. تشير الدراسات إلى أن SGD مع الزخم يمكن أن يحسن سرعة التقارب، بينما أظهرت ADAM وRMSprop أداءً متفوقًا في سياقات معينة، مثل تصنيف أمراض النباتات واكتشاف الآفات. تعتبر مقاييس التقييم مثل الدقة، والاسترجاع، ودرجة F1، والمساحة تحت المنحنى (AUC) ضرورية لتقييم أداء النموذج، خاصة في مجموعات البيانات غير المتوازنة. كما يتم التأكيد على أهمية إعادة التقييم وإمكانية إعادة التدريب استجابةً للتغيرات في توزيع البيانات أو الظروف البيئية، لضمان فعالية النموذج المستمرة في التطبيقات الواقعية.
مناقشة
في قسم المناقشة من الورقة البحثية، يؤكد المؤلفون على الدور الحاسم لاكتساب البيانات ومعالجتها مسبقًا في التعلم الآلي (ML) لتحليل صور النباتات. توفر تقنيات التصوير عالي الدقة، بما في ذلك تصوير الطائرات بدون طيار والمسح ثلاثي الأبعاد، بيانات مورفولوجية مفصلة ضرورية لنماذج التعلم العميق (DL). بينما توفر الأجهزة عالية الدقة جودة متفوقة، تعزز الكاميرات التقليدية الوصول وقابلية التوسع، مما يسهل جمع البيانات على نطاق واسع الذي يحسن تنوع مجموعة البيانات وقوة النموذج. يشمل استخراج الميزات أبعادًا متنوعة مثل اللون، والمورفولوجيا، والقوام، والفسيولوجيا، حيث أثبتت الشبكات العصبية التلافيفية (CNNs) فعاليتها في إدارة الصور النباتية المعقدة لتصنيف واكتشاف الأمراض. كما يبرز المؤلفون أهمية تقنيات المعالجة المسبقة، بما في ذلك القص، وإعادة الحجم، وزيادة البيانات، التي تحسن الصور لنماذج ML ولكن قد تعرض المعلومات للخطر.
تناقش الورقة أيضًا أهمية مجموعات البيانات العامة، مثل مجموعة بيانات قرية النباتات، التي تعتبر معيارًا لتطوير نماذج ML في مراقبة صحة النباتات. يحدد المؤلفون متطلبات حجم مجموعة البيانات المتنوعة بناءً على تعقيد المهمة، مشيرين إلى أن التصنيف الثنائي يتطلب عادةً من 1,000 إلى 2,000 صورة لكل فئة، بينما تتطلب المهام الأكثر تعقيدًا مثل اكتشاف الكائنات مجموعات بيانات أكبر. تغطي المناقشة أيضًا اختيار النموذج لتصنيف الصور واكتشاف الكائنات، حيث تُستخدم الشبكات العصبية التلافيفية (CNNs)، وأقرب الجيران (K-NN)، وآلات الدعم (SVMs) بشكل شائع. تُلاحظ النماذج المتقدمة مثل Faster R-CNN وYOLOv5 لدقتها المكانية العالية في مهام اكتشاف الكائنات. يتم تسليط الضوء على دمج نماذج اللغة (LLMs) في علوم النبات كنهج تحويلي، مما يعزز توقع وظيفة الجين وتشخيص الأمراض من خلال تحليل البيانات متعددة الوسائط. بشكل عام، تؤكد هذه القسم على التقدم في تطبيقات ML ضمن أبحاث النباتات، خاصة في إدارة الإجهاد، وتحديد الأنواع، ومحاكاة النمو، وتقسيم الخلايا، مما يظهر الإمكانات لتحسين الممارسات الزراعية والحلول المستدامة.
DOI: https://doi.org/10.1017/qpb.2025.10018
PMID: https://pubmed.ncbi.nlm.nih.gov/40735612
Publication Date: 2025-01-01
Author(s): Kaiyue Hong et al.
Primary Topic: Smart Agriculture and AI
Overview
Recent advancements in data science and artificial intelligence have significantly impacted plant sciences, particularly through the application of image recognition and deep learning technologies. This review highlights the transformative effects of these innovations on species identification, disease detection, cellular signaling analysis, and growth monitoring. It emphasizes the critical role of data acquisition and preprocessing, detailing techniques such as high-resolution imaging and unmanned aerial vehicle (UAV) photography, as well as image enhancement methods like cropping and scaling.
Furthermore, the review addresses feature extraction techniques, including color histograms and texture analysis, which are vital for effective plant identification and health assessment. The authors also explore emerging trends, challenges, and future directions in the field, providing insights into how these computational tools can advance both plant science research and practical applications.
Introduction
The introduction of the paper emphasizes the critical role of large-scale plant image datasets in advancing plant science, while also addressing the challenges associated with their efficient processing. It highlights the transformative potential of artificial intelligence (AI) and deep learning (DL) in automating plant image analysis through machine learning (ML) techniques. These technologies enable computers to learn from data, identify patterns, and make predictions, significantly enhancing tasks such as plant growth monitoring and disease detection. The paper notes the difficulties posed by plant diversity and complex backgrounds in image processing, which necessitate tailored preprocessing methods, feature extraction, and data augmentation strategies to improve model performance.
Furthermore, the introduction discusses the advancements in image recognition facilitated by Convolutional Neural Networks (CNNs), which have demonstrated high accuracy in plant species recognition across various datasets, outperforming traditional methods. It also introduces the potential of large language models (LLMs) like ChatGPT and the Agronomic Nucleotide Transformer (AgroNT) in plant genetics, showcasing their ability to analyze genomic data and uncover novel insights into gene expression and stress responses. The review aims to evaluate key technologies in plant image processing, comparing methodologies and ML models to provide a comprehensive framework for researchers in the field.
Methods
The methods section outlines the critical components for effective machine learning (ML) model training and evaluation, emphasizing data partitioning, loss functions, and optimization strategies. A standard data split of 70:15:15 for training, validation, and testing is recommended to ensure robust model generalization. Notable studies, such as those by Mohanty et al. and Ferentinos, demonstrate the efficacy of deep learning in plant disease identification and species classification, achieving accuracies exceeding 99%. The section also discusses various loss functions tailored for regression and classification tasks, highlighting the importance of custom loss functions, as exemplified by Gillespie et al.’s work on monitoring rare plant species.
Optimization algorithms play a pivotal role in enhancing model performance, with common techniques including stochastic gradient descent (SGD), adaptive moment estimation (ADAM), and RMSprop. Studies indicate that SGD with momentum can improve convergence speed, while ADAM and RMSprop have shown superior performance in specific contexts, such as plant disease classification and pest detection. Evaluation metrics like accuracy, recall, F1 score, and area under the curve (AUC) are essential for assessing model performance, particularly in imbalanced datasets. The importance of re-evaluation and potential retraining in response to changes in data distribution or environmental conditions is also emphasized, ensuring sustained model efficacy in real-world applications.
Discussion
In the discussion section of the research paper, the authors emphasize the critical role of data acquisition and preprocessing in machine learning (ML) for plant image analysis. High-resolution imaging techniques, including unmanned aerial vehicle (UAV) photography and 3D scanning, provide detailed morphological data essential for deep learning (DL) models. While high-resolution devices yield superior quality, conventional cameras enhance accessibility and scalability, facilitating large-scale data collection that improves dataset diversity and model robustness. Feature extraction encompasses various dimensions such as color, morphology, texture, and physiology, with convolutional neural networks (CNNs) proving effective in managing complex plant images for disease classification and detection. The authors also highlight the importance of preprocessing techniques, including cropping, resizing, and data augmentation, which optimize images for ML models but may risk information loss.
The paper further discusses the significance of public datasets, such as the Plant Village dataset, which serves as a benchmark for developing ML models in plant health monitoring. The authors outline the varying dataset size requirements based on task complexity, noting that binary classification typically requires 1,000 to 2,000 images per class, while more complex tasks like object detection necessitate larger datasets. The discussion also covers model selection for image classification and object detection, with CNNs, K-nearest neighbors (K-NN), and support vector machines (SVMs) being commonly employed. Advanced models like Faster R-CNN and YOLOv5 are noted for their high spatial accuracy in object detection tasks. The integration of language models (LLMs) into plant science is highlighted as a transformative approach, enhancing gene function prediction and disease diagnostics through multimodal data analysis. Overall, the section underscores the advancements in ML applications within plant research, particularly in stress management, species identification, growth simulation, and cell segmentation, showcasing the potential for improved agricultural practices and sustainable solutions.
