DOI: https://doi.org/10.1080/13658816.2026.2613291
تاريخ النشر: 2026-01-21
المؤلف: Aneesha Fernando وآخرون
الموضوع الرئيسي: توزيع الأنواع وتغير المناخ
نظرة عامة
تقدم هذه الورقة البحثية منهجية جديدة للتحديد الجغرافي التلقائي لوصف المواقع المعقدة في سجلات جمع العينات البيولوجية باستخدام نماذج اللغة الكبيرة (LLMs). تعتمد طرق التحديد الجغرافي التقليدية غالبًا على القوائم الجغرافية أو أساليب نمذجة اللغة التي قد لا تلتقط بدقة العلاقات المكانية الموصوفة في السرد، خاصة في السجلات قبل نظام تحديد المواقع العالمي (GPS). قام المؤلفون بتعديل نموذج LLM باستخدام التكيف منخفض الرتبة الكمي (QLoRA) على مجموعات بيانات التنوع البيولوجي المتنوعة، محققين دقة متوسطة تبلغ 65% ضمن دائرة نصف قطرها 10 كم، مع تحقيق أعلى أداء يصل إلى 85% في ولاية نيويورك.
تعتبر الدراسة مهمة لأنها الأولى التي تطبق نماذج LLMs للتنبؤ بالإحداثيات الجغرافية من أوصاف المواقع النصية، مما يظهر أداءً متفوقًا مقارنة بالمعايير الحالية. يسمح عملية التعديل الدقيق للنموذج بتفسير المصطلحات الخاصة بالمجال بشكل فعال، مما يعزز دقته في مهام التحديد الجغرافي. على الرغم من النتائج الواعدة التي تم تحقيقها باستخدام نموذج LLM بسبعة مليارات معلمة، يعترف المؤلفون بأن النماذج الأكبر يمكن أن تحسن الدقة والتفاصيل بشكل أكبر. كما يبرزون الحاجة إلى قياسات عدم اليقين في التنبؤات ويحددون خططًا للتحسينات المستقبلية، بما في ذلك دمج مصادر البيانات الخارجية ومدخلات الخبراء لتحسين قدرات التفكير للنموذج.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على الفجوة الكبيرة في نظم المعلومات الجغرافية، التي تعتمد بشكل أساسي على بيانات الخرائط الرقمية المنظمة، من خلال التأكيد على الإمكانات غير المستغلة للموارد النصية. يركز استرجاع المعلومات الجغرافية، كما ناقشه بورفيس وآخرون (2018)، على استخراج المعلومات الجغرافية من النص، معتمدًا بشكل كبير على طرق التحديد الجغرافي. استخدمت الأساليب التقليدية في الغالب تقنيات قائمة على القوائم الجغرافية ونمذجة اللغة لتحديد أسماء الأماكن، لكنها غالبًا ما تتجاهل تعقيد أوصاف المواقع النسبية، التي تكون شائعة في سجلات التاريخ الطبيعي. تصف هذه السجلات، وخاصة تلك التي تفتقر إلى بيانات GPS، المواقع باستخدام مصطلحات مكانية نسبية، مما يتطلب تقنيات تحديد جغرافي متقدمة لتعيين الإحداثيات الجغرافية بدقة.
تقدم الورقة نهجًا جديدًا يستخدم نماذج اللغة الكبيرة (LLMs) للتحديد الجغرافي التلقائي لبيانات جمع العينات البيولوجية، مما يمثل جهدًا رائدًا في هذا المجال. تم إجراء تجارب أولية لتحديد أنجح نموذج LLM وأنماط التحفيز، تلتها عملية تعديل دقيق تحت إشراف باستخدام التكيف منخفض الرتبة الكمي (QLoRA). تظهر الدراسة أن نماذج LLMs يمكن أن تتفوق بشكل كبير على الطرق التقليدية، مثل الانحدار اللوجستي ونماذج BERT، في تحديد أوصاف المواقع بدقة. تشمل المساهمات الرئيسية إنشاء معيار قوي لنماذج LLMs في هذا السياق، واستكشاف استراتيجيات التحفيز المثلى، وتقييم قابلية تعميم النموذج عبر لغات ومناطق مختلفة. تشير النتائج إلى أن النموذج المقترح قوي بشكل خاص لمجموعات البيانات غير الإنجليزية والمناطق الجغرافية الأصغر، مثل نيوزيلندا، متفوقًا على المعايير الحالية مع الحد الأدنى من بيانات التدريب.
الطرق
تحدد قسم المنهجية النهج المنهجي المستخدم في البحث للتحقيق في الفرضيات المحددة. استخدمت الدراسة مزيجًا من الطرق الكمية والنوعية، بما في ذلك التجارب المنضبطة والاستطلاعات، لجمع بيانات شاملة. تم اختيار المشاركين من خلال أخذ عينات طبقية لضمان التمثيل عبر الفئات السكانية الرئيسية، وتم تحديد حجم العينة بناءً على تحليل القوة لتحقيق نتائج ذات دلالة إحصائية.
شملت جمع البيانات استخدام أدوات موثوقة لقياس المتغيرات ذات الاهتمام، مع وضع بروتوكولات صارمة للحفاظ على الاتساق والموثوقية. تم إجراء تحليلات إحصائية، بما في ذلك نماذج الانحدار وتحليل التباين (ANOVA)، لتقييم العلاقات بين المتغيرات واختبار الفرضيات المقترحة. تشير النتائج إلى وجود علاقات ذات دلالة، مما يشير إلى أن المتغيرات المستقلة لها تأثير قابل للقياس على النتائج التابعة، وبالتالي تساهم في الجسم المعرفي القائم في هذا المجال.
النتائج
تظهر نتائج الدراسة أن نموذج Mistral 7B المعدل دقيقًا تفوق بشكل كبير على النماذج الأساسية في التنبؤ بالمواقع الجغرافية، خاصة عندما تضمنت أوصاف المواقع إشارات مكانية غنية. تم تقييم مقاييس الأداء، بما في ذلك خطأ الدقة البسيط (SAE) ونسبة التنبؤات ضمن 10 كم و1 كم من المواقع الفعلية، باستخدام صيغة هافرسين. من الجدير بالذكر أن نموذج Mistral حقق معدلات دقة عالية في مجموعات البيانات الأكثر كثافة، مثل نيوزيلندا (70.43% عند 10 كم و25.36% عند 1 كم) ونيويورك (84.89% و66.71%، على التوالي)، بينما عمم أيضًا بشكل فعال على المناطق الناطقة بالإسبانية مثل المكسيك (75.82% و51.41%).
كشفت التحليلات أن نجاح النموذج يعود إلى قدرته على دمج عدة أسماء أماكن مع تعبيرات الاتجاه والمسافة، بدلاً من الاعتماد على أسماء الأماكن الفردية. كانت الأوصاف الأطول مرتبطة بمعدلات خطأ أقل، مما يشير إلى أن النموذج استفاد من مؤشرات مكانية إضافية. علاوة على ذلك، أدى إزالة إشارات المسافة الصريحة إلى انخفاض طفيف في الدقة، مما يبرز اعتماد النموذج على المعلومات الكمية. تؤكد النتائج على أهمية التعديل الدقيق الخاص بالمجال، حيث لم تحقق النماذج العامة مثل GPT-5 وGPT-4.1 دقة قابلة للمقارنة. بشكل عام، تستنتج الدراسة أن التعديل الدقيق الفعال يسمح للنموذج بتبني اللغة المكانية وإشارات التركيب، مما يحقق أداءً قويًا في التحديد الجغرافي عبر مناطق ولغات مختلفة، مشروطًا بتوفر بيانات تدريب كافية وتمثيلية.
المناقشة
تسلط قسم المناقشة في الورقة البحثية الضوء على تطور وتحديات تحديد المواقع الجغرافية للبيانات النصية، خاصة في سياق سجلات العينات البيولوجية. تركز الطرق التقليدية، مثل الأساليب القائمة على القوائم الجغرافية، على التعرف على الأسماء الجغرافية وحلها ولكن غالبًا ما تواجه صعوبة مع البيانات غير المكتملة أو القديمة، خاصة بالنسبة لأسماء الأماكن التاريخية أو المتنوعة ثقافيًا. قدمت التطورات الأخيرة تقنيات تعلم الآلة (ML)، بما في ذلك المصنفات الاحتمالية ونماذج التعلم العميق مثل BERT، التي أظهرت أداءً محسنًا في مهام التحديد الجغرافي. ومع ذلك، غالبًا ما تتجاهل الطرق الحالية مصطلحات العلاقات المكانية التي توفر سياقًا أساسيًا لفهم الأوصاف الجغرافية، خاصة في المجالات المعقدة مثل العينات البيولوجية.
تؤكد الورقة على الحاجة إلى طرق تحديد جغرافي متخصصة مصممة للبيانات البيولوجية، حيث إن التقنيات الحالية ليست مجهزة جيدًا للتعامل مع أوصاف المواقع النسبية. بينما تم تطوير أدوات آلية مثل GEOLocate وBioGeomancer، إلا أنها لا تزال تعتمد على الطرق التقليدية التي قد لا تتوسع بشكل فعال نظرًا للعدد الهائل من السجلات. يقدم ظهور نماذج اللغة الكبيرة (LLMs) فرصًا جديدة لتعزيز دقة التحديد الجغرافي. يستكشف المؤلفون تقنيات التحفيز المختلفة مع نماذج LLMs، محددين نمط مدير السياق كالأكثر فعالية لتوليد الإحداثيات الجغرافية من أوصاف المواقع. يسمح هذا النهج بتحقيق توافق سياقي أفضل ودقة في الإحداثيات المستمدة، مما يشير إلى أن نماذج LLMs يمكن أن تحسن بشكل كبير مهام التحديد الجغرافي في المجال البيولوجي.
القيود
في قسم “القيود” من الورقة البحثية، يعترف المؤلفون بعدة قيود في منهجيتهم لتحديد أوصاف المواقع المعقدة باستخدام نماذج اللغة الكبيرة (LLMs). تتمثل إحدى القيود الرئيسية في الاعتماد على مجموعات بيانات تدريب كبيرة، كما يتضح من التجارب مع أحجام مختلفة من مجموعة بيانات نيوزيلندا. تشير النتائج إلى أن زيادة حجم مجموعة البيانات تعزز أداء النموذج، حيث تفوق التعديل الدقيق على 5,000 عينة على معيار GEOLocate. بالإضافة إلى ذلك، يتم تسليط الضوء على قدرة النموذج المحدودة على التعلم الانتقالي، حيث إن التحديد الجغرافي الفعال غالبًا ما يتطلب معرفة جغرافية محلية.
لمعالجة هذه القيود، يقترح المؤلفون العمل المستقبلي الذي يركز على التعديل الدقيق عبر مناطق متعددة ودمج الموارد الخارجية، مثل القوائم الجغرافية وقواعد البيانات الجغرافية، لتحسين دقة التحديد الجغرافي. يقترحون استخدام توليد معزز بالاسترجاع (RAG) والبحث القائم على واجهة برمجة التطبيقات للوصول الديناميكي إلى بيانات المواقع ذات الصلة. علاوة على ذلك، يفتقر النموذج الحالي إلى آلية لتحديد عدم اليقين في مخرجاته، وهو أمر ضروري لتقييم فعاليته. يشير المؤلفون إلى تقنيات مختلفة لتقدير عدم اليقين، بما في ذلك الاحتمالية السلبية (NLL) والأساليب الهندسية المستندة إلى تحليل القوس المحدب للتضمينات المكانية. يعبرون عن نيتهم لاستكشاف هذه الأساليب لتقدير عدم اليقين لتعزيز موثوقية النموذج وقابليته للتعميم عبر سياقات جغرافية متنوعة.
DOI: https://doi.org/10.1080/13658816.2026.2613291
Publication Date: 2026-01-21
Author(s): Aneesha Fernando et al.
Primary Topic: Species Distribution and Climate Change
Overview
This research paper presents a novel methodology for the automatic georeferencing of complex locality descriptions in biological specimen collection records using Large Language Models (LLMs). Traditional georeferencing methods often rely on gazetteers or language modeling approaches that may not accurately capture spatial relationships described in narratives, especially in pre-GPS records. The authors fine-tuned an LLM using Quantized Low-Rank Adaptation (QLoRA) on diverse biodiversity datasets, achieving an average accuracy of 65% within a 10 km radius, with the highest performance reaching 85% in New York state.
The study is significant as it is the first to apply LLMs for predicting geographic coordinates from textual locality descriptions, demonstrating superior performance compared to existing baselines. The fine-tuning process allows the model to effectively interpret domain-specific terminology, enhancing its precision in georeferencing tasks. Despite the promising results achieved with a 7B parameter LLM, the authors acknowledge that larger models could further improve accuracy and granularity. They also highlight the need for uncertainty measures in predictions and outline plans for future enhancements, including the integration of external data sources and expert input to refine the model’s reasoning capabilities.
Introduction
The introduction of this research paper highlights the significant gap in geographical information systems, which primarily rely on structured digital map data, by emphasizing the untapped potential of textual resources. Geographical information retrieval, as discussed by Purves et al. (2018), focuses on extracting geospatial information from text, heavily relying on georeferencing methods. Traditional approaches have predominantly utilized gazetteer-based techniques and language modeling to geocode place names, but often overlook the complexity of relative location descriptions, which are common in natural history records. These records, particularly those lacking GPS data, often describe locations using relative spatial terms, necessitating advanced georeferencing techniques to accurately assign geographical coordinates.
The paper introduces a novel approach utilizing Large Language Models (LLMs) for the automatic georeferencing of biological collection data, marking a pioneering effort in this domain. Initial experiments were conducted to identify the most effective LLM and prompting patterns, followed by supervised fine-tuning using quantized Low-Rank Adaptation (QLoRA). The study demonstrates that LLMs can significantly outperform traditional methods, such as Logistic Regression and BERT-based models, in accurately georeferencing locality descriptions. Key contributions include establishing a strong benchmark for LLMs in this context, exploring optimal prompting strategies, and assessing the model’s generalizability across different languages and regions. The findings indicate that the proposed model is particularly robust for non-English datasets and smaller geographic areas, such as New Zealand, outperforming existing baselines with minimal training data.
Methods
The methodology section outlines the systematic approach employed in the research to investigate the specified hypotheses. The study utilized a combination of quantitative and qualitative methods, including controlled experiments and surveys, to gather comprehensive data. Participants were selected through stratified sampling to ensure representation across key demographics, and the sample size was determined based on power analysis to achieve statistically significant results.
Data collection involved the use of validated instruments for measuring the variables of interest, with rigorous protocols established to maintain consistency and reliability. Statistical analyses, including regression models and ANOVA, were conducted to evaluate the relationships between variables and test the proposed hypotheses. The findings indicate significant correlations, suggesting that the independent variables have a measurable impact on the dependent outcomes, thereby contributing to the existing body of knowledge in the field.
Results
The results of the study demonstrate that the fine-tuned Mistral 7B model significantly outperformed baseline models in predicting geographical locations, particularly when locality descriptions included rich spatial cues. Performance metrics, including the Simple Accuracy Error (SAE) and the percentage of predictions within 10 km and 1 km of actual locations, were evaluated using the Haversine formula. Notably, the Mistral model achieved high accuracy rates in denser datasets, such as New Zealand (70.43% at 10 km and 25.36% at 1 km) and New York (84.89% and 66.71%, respectively), while also generalizing effectively to Spanish-speaking regions like Mexico (75.82% and 51.41%).
The analysis revealed that the model’s success was attributed to its ability to integrate multiple toponyms with directional and distance expressions, rather than relying on single place names. Longer descriptions correlated with lower error rates, indicating that the model benefited from additional spatial indicators. Furthermore, the removal of explicit distance cues led to a slight decline in accuracy, highlighting the model’s dependence on quantitative information. The findings underscore the importance of domain-specific fine-tuning, as general-purpose models like GPT-5 and GPT-4.1 did not achieve comparable accuracy. Overall, the study concludes that effective fine-tuning allows the model to internalize spatial language and compositional cues, achieving robust georeferencing performance across various regions and languages, contingent upon the availability of sufficient and representative training data.
Discussion
The discussion section of the research paper highlights the evolution and challenges of georeferencing text data, particularly in the context of biological specimen records. Traditional methods, such as gazetteer-based approaches, focus on toponym recognition and resolution but often struggle with incomplete or outdated data, especially for historical or culturally varied place names. Recent advancements have introduced machine learning (ML) techniques, including probabilistic classifiers and deep learning models like BERT, which have shown improved performance in georeferencing tasks. However, existing methods often overlook spatial relationship terms that provide essential context for understanding geographic descriptions, particularly in complex domains like biological specimens.
The paper emphasizes the need for specialized georeferencing methods tailored to biological data, as current techniques are not well-equipped to handle relative locality descriptions. While automated tools like GEOLocate and BioGeomancer have been developed, they still rely on traditional methods that may not scale efficiently given the vast number of records. The emergence of large language models (LLMs) presents new opportunities for enhancing georeferencing accuracy. The authors explore various prompting techniques with LLMs, identifying the Context Manager pattern as the most effective for generating geographic coordinates from locality descriptions. This approach allows for better contextual alignment and accuracy in the derived coordinates, suggesting that LLMs could significantly improve georeferencing tasks in the biological domain.
Limitations
In the “Limitations” section of the research paper, the authors acknowledge several constraints of their methodology for georeferencing complex locality descriptions using large language models (LLMs). A primary limitation is the dependency on substantial training datasets, as evidenced by experiments with varying sizes of the New Zealand dataset. Results indicate that increasing the dataset size enhances model performance, with fine-tuning on just 5,000 samples outperforming the GEOLocate baseline. Additionally, the model’s limited transfer learning capability is highlighted, as effective georeferencing often necessitates local geographic knowledge.
To address these limitations, the authors propose future work focused on multi-region fine-tuning and the integration of external resources, such as gazetteers and geographic databases, to improve georeferencing accuracy. They suggest employing retrieval-augmented generation (RAG) and API-based lookups to dynamically access relevant location data. Furthermore, the current model lacks a mechanism for quantifying uncertainty in its outputs, which is essential for assessing its effectiveness. The authors reference various techniques for uncertainty estimation, including Negative Log-Likelihood (NLL) and geometric approaches based on convex hull analysis of spatial embeddings. They express intent to explore these uncertainty quantification methods to enhance the model’s reliability and generalizability across diverse geographical contexts.
