ما مدى فعالية نماذج الأساس في رسم خرائط أنواع المحاصيل باستخدام التصوير الطيفي عالي الدقة؟ دراسة مقارنة بين التعلم الآلي، التعلم العميق، ونماذج الأساس الجغرافية
How Effective Are Foundation Models for Crop Type Mapping Using Hyperspectral Imaging? A Comparative Study of Machine Learning, Deep Learning, and Geospatial Foundation Models

شارك:
المجلة: ˜The œinternational archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences
DOI: https://doi.org/10.5194/isprs-archives-xlviii-4-w17-2025-69-2026
تاريخ النشر: 2026-01-15
المؤلف: Mohamed Bourriz وآخرون
الموضوع الرئيسي: الاستشعار عن بعد في الزراعة

نظرة عامة

تقيّم هذه الدراسة أداء خوارزميات التعلم الآلي التقليدية (ML) ونماذج التعلم العميق (DL) ونماذج الأساس الجغرافية (GFMs) لتصنيف أنواع المحاصيل الطيفية، باستخدام مجموعة بيانات معيارية من الهند. تتناول الأبحاث التحديات المتعلقة بنقص البيانات في مراقبة الزراعة، مع تسليط الضوء على القدرات الفائقة للنماذج التي تتعلم بفعالية الميزات المكانية والطيفية. ومن الجدير بالذكر أن HybridSN حقق دقة قريبة من الكمال في السيناريوهات ذات العلامات العالية، بينما تفوق HyperSIGMA في الظروف ذات العلامات المنخفضة، مما يظهر مزايا التعلم الذاتي. أدت تكييف نموذج الأساس متعدد الطيف Prithvi-EO-1.0 مع البيانات الطيفية العالية إلى نتائج واعدة، حيث حقق دقة عامة (OA) تصل إلى 97%، على الرغم من تراجع أدائه في البيئات ذات العلامات المنخفضة.

تشير النتائج إلى أن نماذج ML التقليدية، مثل آلات الدعم الناقل (SVM) والغابات العشوائية (RF)، تكافح لالتقاط التفاعلات الطيفية-المكانية المعقدة والتعميم بفعالية تحت قيود العلامات. تقترح الدراسة أن التعديلات المستقبلية، مثل RF المكاني أو XG-Boost المكاني، يمكن أن تعزز أداء النماذج التقليدية من خلال دمج السياق المحلي. تؤكد الأبحاث على إمكانية تطبيق هذه النماذج المتقدمة على الصور الطيفية العالية من الفضاء، مما يسهل رسم الخرائط الزراعية على نطاق واسع ويساهم في الأمن الغذائي وإدارة الأراضي المستدامة.

مقدمة

تتناول مقدمة هذه الورقة البحثية القضية الحرجة للأمن الغذائي وضرورة الحصول على بيانات دقيقة لتوزيع المحاصيل الجغرافية، وهو أمر حيوي لإدارة الأراضي الفعالة، ومراقبة المحاصيل، وتقدير العائد، والحفاظ على التربة. لقد ظهرت الاستشعار عن بعد كتقنية رئيسية لرسم الخرائط الزراعية على نطاق واسع، حيث يتم استخدام الصور متعددة الطيف (MSI) بشكل شائع جنبًا إلى جنب مع طرق التعلم الآلي (ML) والتعلم العميق (DL). ومع ذلك، غالبًا ما تكافح MSI بسبب نطاقاتها الطيفية المحدودة لتمييز بين أنواع المحاصيل المتشابهة، خاصة في مراحل النمو المبكرة أو أنظمة الزراعة المختلطة. بالمقابل، توفر الصور الطيفية العالية (HSI) دقة طيفية عالية، مما يمكّن من تمييز المحاصيل بشكل أفضل من خلال معلومات طيفية مفصلة.

تسلط الورقة الضوء على قيود خوارزميات ML التقليدية، مثل آلات الدعم الناقل (SVMs) والغابات العشوائية (RFs)، التي تعتمد بشكل أساسي على البيانات الطيفية وتتجاهل السياق المكاني، مما يقلل من فعاليتها في المناظر الطبيعية غير المتجانسة. لمعالجة هذه التحديات، أظهرت طرق DL، وخاصة الشبكات العصبية التلافيفية (CNNs) والهياكل الهجينة مثل شبكة Hybrid Spectral-Spatial (HybridSN) وCVT-Net، وعدًا في التقاط كل من الميزات الطيفية والمكانية. لقد أظهرت هذه النماذج أداءً متفوقًا على مجموعات البيانات المعيارية، محققة تحسينات كبيرة في دقة التصنيف مقارنة بالطرق التقليدية. ومع ذلك، فإن اعتمادها على مجموعات بيانات موسومة كبيرة وموارد حسابية كبيرة يمثل تحديات. يقدم إدخال نماذج الأساس الجغرافية (GFMs)، التي تستفيد من التعلم الذاتي (SSL) على مجموعات بيانات غير موسومة واسعة، حلاً محتملاً من خلال تقليل الحاجة إلى البيانات الموسومة. تهدف هذه الدراسة إلى تقييم أداء ML التقليدي ونماذج DL وGFMs لتصنيف المحاصيل، مع التركيز على تكييف نموذج Prithvi-EO 1.0 المدرب مسبقًا مع المجال الطيفي العالي، وبالتالي استكشاف جدوى التعلم الانتقالي من متعدد الطيف إلى الطيف العالي في السيناريوهات التي تعاني من نقص البيانات.

الطرق

في هذه الدراسة، تم هيكلة المنهجية في ثلاث مراحل رئيسية: إعداد البيانات، تقييم النموذج، وتحليل الأداء، باستخدام مجموعة بيانات الهند الطيفية العالية. تم تقييم خمسة نماذج تصنيف، مصنفة إلى ثلاث طرق منهجية: تقنيات التعلم الآلي التقليدية (ML) (آلة الدعم الناقل (SVM) والغابة العشوائية (RF))، التعلم العميق (شبكة عصبية تلافيفية 1D (1D-CNN) وHybridSN)، والأطر القائمة على الرسوم البيانية (GFMs) (Hyper-SIGMA وPrithvi-EO 1.0). خضعت كل نموذج لتدريب وتقييم مستقل عبر أربعة سيناريوهات تجريبية، والتي اختلفت في توفر بيانات التدريب، كما هو موضح في الجدول 2.

شملت السيناريوهات التجريبية تقسيمات مختلفة من مجموعات بيانات التدريب والتحقق والاختبار، تتراوح من 70% من بيانات التدريب في السيناريو 1 إلى 1% فقط في السيناريو 4. تم استخدام مقاييس الأداء مثل الدقة العامة (OA) ومعامل كابا لتقييم فعالية النموذج. تم اختبار النماذج التي حققت أفضل أداء في السيناريوهات 1 إلى 3 تحت ظروف العلامات المنخفضة القصوى في السيناريو 4، بهدف تحديد جدواها في البيئات التي تعاني من نقص البيانات. يوفر هذا التقييم الشامل رؤى حول قوة طرق التصنيف المختلفة لتصنيف المحاصيل الطيفية العالية.

النتائج

قيمت نتائج الدراسة نماذج مختلفة لتصنيف الصور الطيفية العالية تحت أربعة سيناريوهات (S1-S4) التي تحاكي مستويات مختلفة من توفر البيانات الموسومة. تفوقت نموذج HybridSN باستمرار على الآخرين في السيناريوهات S1 إلى S3، محققة قيم دقة عامة (OA) تتجاوز 98.9% ومعامل كابا فوق 0.98، وذلك بفضل هيكله الهجين الذي يلتقط بفعالية الميزات الطيفية-المكانية من خلال مزيج من التلافيف ثلاثية الأبعاد وثنائية الأبعاد. ومع ذلك، تراجع أداؤه بشكل كبير في السيناريو المنخفض العلامات S4 (OA = 65.75%)، مما يشير إلى اعتماده على توفر بيانات موسومة كافية. بالمقابل، أظهر HyperSIGMA أداءً قويًا عبر جميع السيناريوهات، مستفيدًا من هيكل المحول وآليات الانتباه الذاتي لنمذجة الاعتماديات بعيدة المدى، متفوقًا بشكل خاص في الظروف ذات العلامات المنخفضة بفضل تدريبه الذاتي.

كان نموذج Prithvi-EO 1.0، على الرغم من فعاليته في السيناريوهات S1 وS2، يعاني في S3 وS4 بسبب نقص الغنى الطيفي وعدم وجود ضبط دقيق. حققت نماذج التعلم الآلي التقليدية مثل الغابة العشوائية (RF) وآلة الدعم الناقل (SVM) أداءً معتدلاً في السيناريوهات ذات العلامات العالية ولكن أظهرت تدهورًا في S3، مما يبرز قدراتها المحدودة على التعميم. تقترح الدراسة أن التحسينات المستقبلية يمكن أن تشمل دمج آليات الانتباه في نماذج مثل HybridSN وتطبيق تقنيات الضبط الدقيق مثل التكيف منخفض الرتبة (LoRA) لنماذج مثل HyperSIGMA وPrithvi-EO 1.0. بالإضافة إلى ذلك، يمكن أن يؤدي دمج الميزات المكانية في نماذج ML التقليدية إلى تحسين أدائها في حل الغموض في الفئات الطيفية المتشابهة. تؤكد الأبحاث على الحاجة إلى مزيد من الاستكشاف باستخدام مجموعات بيانات طيفية عالية النطاق من المنصات الفضائية لتعزيز التطبيق العملي للمنهجيات المقترحة في السياقات الزراعية.

المناقشة

تركز قسم المناقشة في الدراسة على تقييم نماذج التصنيف المختلفة لتصنيف نوع المحاصيل الطيفية العالية باستخدام مجموعة بيانات الهند. تتضمن هذه المجموعة، التي تم الحصول عليها بواسطة مستشعر AVIRIS، 220 نطاقًا طيفيًا وتنوعًا في تغطية الأراضي الزراعية، مما يجعلها معيارًا لاختبار خوارزميات التعلم الآلي (ML) والتعلم العميق (DL) ونماذج الأساس الجغرافية (GFMs). استخدمت الدراسة مصنفات تقليدية مثل آلات الدعم الناقل (SVM) والغابات العشوائية (RF)، التي تم تحسينها للمعلمات الفائقة ولكنها لم تتضمن المعلومات المكانية. بالمقابل، تم تصميم نماذج DL مثل الشبكة العصبية التلافيفية ذات البعد الواحد (1D-CNN) وشبكة Hybrid Spectral-Spatial (HybridSN) للاستفادة من كل من الميزات الطيفية والمكانية، حيث أظهرت HybridSN أداءً متفوقًا بفضل قدراتها في استخراج الميزات الهرمية.

شملت GFMs التي تم تقييمها HyperSIGMA، وهو محول رؤية تم تدريبه مسبقًا على مجموعة بيانات طيفية عالية كبيرة، وإصدار معدّل من Prithvi-EO 1.0، الذي تم تعديله للبيانات الطيفية العالية من خلال المحاكاة الطيفية. أشارت النتائج إلى أن النماذج القادرة على دمج المعلومات المكانية والطيفية، وخاصة HybridSN وHyperSIGMA، تفوقت بشكل كبير على الطرق التقليدية، خاصة في السيناريوهات ذات العلامات المنخفضة. بينما حققت HybridSN دقة عالية في الإعدادات ذات العلامات الجيدة، تفوق HyperSIGMA في الظروف التي تعاني من نقص البيانات، مما يبرز فعالية التدريب الذاتي. كما لاحظت الدراسة قيود نماذج ML التقليدية في التقاط التفاعلات المعقدة واقترحت تعديلات مستقبلية لتعزيز أدائها في مراقبة الزراعة التشغيلية، خاصة مع تزايد الوصول إلى الصور الطيفية العالية الدقة للتطبيقات على نطاق واسع.

Journal: ˜The œinternational archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences
DOI: https://doi.org/10.5194/isprs-archives-xlviii-4-w17-2025-69-2026
Publication Date: 2026-01-15
Author(s): Mohamed Bourriz et al.
Primary Topic: Remote Sensing in Agriculture

Overview

This study evaluates the performance of traditional machine learning (ML) algorithms, deep learning (DL) models, and geospatial foundation models (GFMs) for hyperspectral crop type classification, utilizing the Indian Pines benchmark dataset. The research addresses the challenges of data scarcity in agricultural monitoring, highlighting the superior capabilities of models that effectively learn spatial and spectral features. Notably, HybridSN achieved near-perfect accuracy in high-label scenarios, while HyperSIGMA excelled in low-label conditions, demonstrating the advantages of self-supervised learning. The adaptation of the multi-spectral foundation model Prithvi-EO-1.0 to hyperspectral data yielded promising results, achieving an overall accuracy (OA) of up to 97%, although its performance declined in low-label environments.

The findings indicate that classical ML models, such as Support Vector Machines (SVM) and Random Forests (RF), struggle to capture complex spectral-spatial interactions and generalize effectively under label constraints. The study suggests that future adaptations, like Spatial RF or Spatial XG-Boost, could enhance the performance of traditional models by integrating local context. The research underscores the potential of applying these advanced models to spaceborne hyperspectral imagery, facilitating large-scale agricultural mapping and contributing to food security and sustainable land management.

Introduction

The introduction of this research paper addresses the critical issue of food security and the necessity for accurate geographical crop distribution data, which is vital for effective land management, crop monitoring, yield estimation, and soil conservation. Remote sensing has emerged as a key technique for large-scale agricultural mapping, with multi-spectral imagery (MSI) being commonly utilized alongside machine learning (ML) and deep learning (DL) methods. However, MSI’s limited spectral bands often struggle to differentiate between similar crop types, particularly in early growth stages or mixed cropping systems. In contrast, hyperspectral imaging (HSI) offers high spectral resolution, enabling better crop discrimination through detailed spectral information.

The paper highlights the limitations of traditional ML algorithms, such as Support Vector Machines (SVMs) and Random Forests (RFs), which primarily rely on spectral data and ignore spatial context, thus reducing their effectiveness in heterogeneous landscapes. To address these challenges, DL approaches, particularly Convolutional Neural Networks (CNNs) and hybrid architectures like Hybrid Spectral-Spatial Network (HybridSN) and CVT-Net, have shown promise in capturing both spectral and spatial features. These models have demonstrated superior performance on benchmark datasets, achieving significant improvements in classification accuracy compared to traditional methods. However, their reliance on large labeled datasets and substantial computational resources poses challenges. The introduction of geospatial foundation models (GFMs), which leverage self-supervised learning (SSL) on extensive unlabeled datasets, presents a potential solution by reducing the need for labeled data. This study aims to evaluate the performance of classical ML, DL models, and GFMs for crop classification, focusing on the adaptation of the multispectral pretrained model Prithvi-EO 1.0 to the hyperspectral domain, thereby exploring the feasibility of multispectral-to-hyperspectral transfer learning in data-scarce scenarios.

Methods

In this study, the methodology is structured into three key stages: data preparation, model evaluation, and performance analysis, utilizing the Indian Pines hyperspectral dataset. Five classification models were assessed, categorized into three methodological approaches: traditional machine learning (ML) techniques (Support Vector Machine (SVM) and Random Forest (RF)), deep learning (1D Convolutional Neural Network (1D-CNN) and HybridSN), and graph-based frameworks (GFMs) (Hyper-SIGMA and Prithvi-EO 1.0). Each model underwent independent training and evaluation across four experimental scenarios, which varied in training data availability, as detailed in Table 2.

The experimental scenarios included different splits of training, validation, and testing datasets, ranging from 70% training data in Scenario 1 to only 1% in Scenario 4. Performance metrics such as Overall Accuracy (OA) and the Kappa coefficient were employed to evaluate model effectiveness. The models that performed best in Scenarios 1 to 3 were further tested under the extreme low-label conditions of Scenario 4, aiming to determine their viability in data-scarce environments. This comprehensive evaluation provides insights into the robustness of various classification methods for hyperspectral crop classification.

Results

The results of the study evaluated various models for hyperspectral image classification under four scenarios (S1-S4) that simulated different levels of labeled data availability. The HybridSN model consistently outperformed others in scenarios S1 to S3, achieving overall accuracy (OA) values exceeding 98.9% and Kappa coefficients above 0.98, attributed to its hybrid architecture that effectively captures spectral-spatial features through a combination of 3D and 2D convolutions. However, its performance declined significantly in the low-label scenario S4 (OA = 65.75%), indicating a dependency on sufficient labeled data. In contrast, HyperSIGMA demonstrated robust performance across all scenarios, leveraging a transformer backbone and self-attention mechanisms to model long-range dependencies, particularly excelling in low-label conditions due to its self-supervised pretraining.

The Prithvi-EO 1.0 model, while effective in scenarios S1 and S2, suffered in S3 and S4 due to limited spectral richness and lack of fine-tuning. Traditional machine learning models like Random Forest (RF) and Support Vector Machine (SVM) performed moderately in higher-label scenarios but showed degradation in S3, highlighting their limited generalization capabilities. The study suggests that future enhancements could include integrating attention mechanisms in models like HybridSN and employing fine-tuning techniques such as Low-Rank Adaptation (LoRA) for models like HyperSIGMA and Prithvi-EO 1.0. Additionally, incorporating spatial features in classical ML models could improve their performance in resolving ambiguities in spectrally similar classes. The research emphasizes the need for further exploration using large-scale hyperspectral datasets from space-borne platforms to enhance the practical applicability of the proposed methodologies in agricultural contexts.

Discussion

The discussion section of the study focuses on the evaluation of various classification models for hyperspectral crop type classification using the Indian Pines dataset. This dataset, acquired by the AVIRIS sensor, includes 220 spectral bands and diverse agricultural land covers, making it a benchmark for testing machine learning (ML), deep learning (DL), and geospatial foundation models (GFMs). The study employed traditional classifiers such as Support Vector Machines (SVM) and Random Forests (RF), which were optimized for hyperparameters but did not incorporate spatial information. In contrast, DL models like the one-dimensional convolutional neural network (1D-CNN) and the Hybrid Spectral-Spatial Network (HybridSN) were designed to leverage both spectral and spatial features, with HybridSN demonstrating superior performance due to its hierarchical feature extraction capabilities.

The GFMs evaluated included HyperSIGMA, a vision transformer pretrained on a large hyperspectral dataset, and an adapted version of Prithvi-EO 1.0, which was modified for hyperspectral data through spectral simulation. The results indicated that models capable of integrating spatial and spectral information, particularly HybridSN and HyperSIGMA, significantly outperformed traditional methods, especially in low-label scenarios. While HybridSN achieved high accuracy in well-labeled settings, HyperSIGMA excelled in data-scarce conditions, highlighting the effectiveness of self-supervised pretraining. The study also noted the limitations of classical ML models in capturing complex interactions and suggested future adaptations to enhance their performance in operational agricultural monitoring, particularly as high-resolution hyperspectral imagery becomes more accessible for large-scale applications.

شارك: