DOI: https://doi.org/10.1017/pan.2025.10017
تاريخ النشر: 2025-09-19
المؤلف: Andrew Halterman وآخرون
الموضوع الرئيسي: الذكاء الاصطناعي في القانون
نظرة عامة
يتناول هذا القسم من ورقة البحث دمج نماذج اللغة الكبيرة التوليدية (LLMs) في ترميز النصوص السياسية، مع التأكيد على الحاجة إلى التحقق التجريبي من أدائها مقابل العمليات التشغيلية المعتمدة في كتب الرموز. يقدم المؤلفون إطار عمل من خمس مراحل لتقييم LLMs في هذا السياق، والذي يتضمن إعداد كتب الرموز، واختبار قدرات LLM، وتقييم الدقة بدون عينة، وتحليل الأخطاء، وإجراء التدريب الخاضع للإشراف. يظهرون تجريبياً هذا الإطار باستخدام ثلاثة كتب رموز في العلوم السياسية تتعلق بأحداث الاحتجاج، والعنف السياسي، والبيانات الانتخابية، مما يكشف أنه بينما تكافح نماذج LLMs ذات الأوزان المفتوحة الحالية مع الالتزام بدون عينة بكتب الرموز، يمكن أن يعزز التدريب الخاضع للإشراف أدائها بشكل كبير.
في الختام، يبرز المؤلفون التحديات المميزة لاستخدام LLMs في القياس القائم على كتب الرموز، مما يتطلب الالتزام بمعايير ترميز مفصلة بدلاً من الاعتماد فقط على التدريب المسبق. يدعون إلى استخدام إطار العمل المكون من خمس مراحل لتوجيه الباحثين في تقييم LLMs لمهام القياس المحددة، مشيرين إلى أنه مع تقدم تكنولوجيا LLM، من المتوقع أن تتحسن دقة القياس. تسهم الورقة في توفير مجموعات بيانات قيمة وإرشادات منهجية، مما يمكّن الباحثين من اتخاذ قرارات مستنيرة بشأن إعداد كتب الرموز، واختيار النماذج، والاختيار بين الأساليب بدون عينة والخاضعة للإشراف.
مقدمة
في السنوات الأخيرة، استخدم علماء السياسة بشكل متزايد نماذج اللغة الكبيرة التوليدية (LLMs) لتحليل النصوص، خاصةً لقدرتها على تصنيف الوثائق بطريقة “بدون عينة” دون أمثلة مصنفة. تتضمن هذه الطريقة تزويد LLMs بتسميات فئوية ووصف موجز لتوليد التنبؤات لتصنيفات الوثائق. ومع ذلك، تثير هذه الطريقة مخاوف بشأن صلاحية القياس في العلوم السياسية، حيث تفترض أن التعليمات البسيطة يمكن أن تقوم بتشغيل مفاهيم معقدة بدقة. يبرز المؤلفون أن LLMs، المدربة على مجموعات بيانات واسعة، قد تظهر دقة تنبؤية معقولة للمفاهيم المحددة بشكل عام ولكن يمكن أن تسيء تفسير التعريفات المحددة، مما يؤدي إلى عدم دقة محتملة في التصنيف.
لمعالجة هذه القضايا، يقترح المؤلفون إطار تقييم مكون من خمس مراحل يهدف إلى تعزيز قياس كتب الرموز-LLM. يتضمن هذا الإطار مراحل لإعداد كتب الرموز، واختبار بدون تسميات، وتقييم بدون عينة، وتقييمات القوة، وتدريب LLMs الخاضع للإشراف. يقدمون سبعة اختبارات سلوكية جديدة محددة لكتب الرموز يمكن تطبيقها في وقت مبكر من مشاريع البحث دون الحاجة إلى تسميات الحقيقة الأساسية. بالإضافة إلى ذلك، يقدم المؤلفون ثلاثة كتب رموز في العلوم السياسية تتعلق بالسياسة المثيرة للجدل، والعنف السياسي، والبيانات الانتخابية، والتي تشكل تحديات لـ LLMs بسبب تعقيدها. تكشف نتائجهم التجريبية عن قيود في التزام LLMs بتعليمات كتب الرموز في سيناريوهات بدون عينة، لكنهم يظهرون أن التدريب الخاضع للإشراف يمكن أن يعزز الأداء بشكل كبير، مما يحسن دقة التصنيف بنسبة تصل إلى 55%. تختتم الورقة بإرشادات للباحثين حول تنفيذ هذه التقنيات التدريبية بشكل فعال.
النتائج
في هذا القسم، تم تقييم أداء أربعة نماذج لغة كبيرة (LLMs)—Mistral-7B، Mistral-NeMo-12B، وLlama-8B—باستخدام اختبارات سلوكية تعتمد على كتاب الرموز BFRS. أظهرت النماذج أداءً قويًا في الاختبارات من I إلى III، والتي قيمت المخرجات القانونية واستعادة التعريفات والأمثلة. ومع ذلك، أظهرت جميع النماذج حساسية لترتيب أوصاف الفئات في كتاب الرموز، مما يشير إلى وجود مشكلات محتملة في الانتباه عند معالجة المطالبات الطويلة. كشفت التقييمات بدون عينة عن درجات F1 منخفضة مرجحة، حيث حققت Mistral-7B وLlama-8B درجات قدرها 0.21 في البيانات الانتخابية ودرجات هامشية قدرها 0.65 و0.57 في CCC وBFRS، على التوالي. يشير هذا إلى أن LLMs تكافح مع المهام المعقدة لكتب الرموز، مما يتطلب تحسينات في قدرتها على فهم واتباع التعليمات.
أشارت التحليلات الإضافية من خلال تجارب الإزالة إلى أن الإزالة المنهجية للمكونات أدت إلى انخفاض الأداء، مما يبرز أهمية هذه المكونات لتحقيق درجات F1 الأساسية. من الجدير بالذكر أن حذف تعريفات الفئات أحيانًا حسّن الأداء، مما يشير إلى أن النماذج قد لا تلتزم تمامًا بتعريفات التسميات. كشفت التحليلات اليدوية أنه بينما أظهرت Mistral-7B امتثالًا عاليًا لتعليمات المطالبات لـ BFRS وCCC، إلا أنها أظهرت عدم امتثال كبير للبيانات الانتخابية، بما في ذلك الهلاوس وتنبؤات التسميات غير الصحيحة بناءً على التداخل اللفظي بدلاً من الفهم الدقيق لتعريفات كتب الرموز. أدت جهود ضبط التعليمات إلى تحسينات في الأداء، خاصةً بالنسبة لـ Mistral-7B، التي شهدت زيادة نسبتها 55% في درجة F1 على BFRS بعد الضبط الدقيق. تؤكد هذه النتائج على إمكانية تحسين أداء LLM في المهام المعقدة لكتب الرموز من خلال ضبط التعليمات، على الرغم من زيادة المتطلبات الحاسوبية.
المناقشة
في هذا القسم، يقدم المؤلفون إطارًا شاملاً لتقييم قياس كتب الرموز-LLM، مع التركيز على التصنيف متعدد الفئات للوثائق باستخدام نماذج اللغة الكبيرة (LLMs) الموجهة بواسطة كتاب رموز مكتوب بواسطة البشر. يتكون الإطار من خمس مراحل: (1) **إعداد كتاب الرموز**، حيث يقوم الباحثون بتشغيل مفاهيم العلوم الاجتماعية في تنسيق شبه منظم؛ (2) **اختبار سلوكي بدون تسميات**، الذي يقيم قدرة LLM على اتباع التعليمات دون الحاجة إلى بيانات مصنفة؛ (3) **تقييم بدون عينة مع تسميات**، حيث يتم استخدام مجموعة بيانات صغيرة مصنفة لتقييم أداء LLM؛ (4) **تحليل الأخطاء بدون عينة**، الذي يتضمن تحليل الأخطاء المنهجي ودراسات الإزالة لفهم نقاط ضعف LLM؛ و(5) **الضبط الدقيق الخاضع للإشراف**، حيث قد يختار الباحثون ضبط LLM إذا كان الأداء بدون عينة غير كافٍ.
يبرز المؤلفون أنه بينما أظهرت الدراسات السابقة إمكانية LLMs في تصنيف مفاهيم العلوم الاجتماعية، فإن إطارهم يؤكد على أهمية التشغيل الدقيق وكتب الرموز المنظمة لتحسين دقة التصنيف. يشيرون إلى أنه حتى LLMs ذات الأداء العالي يمكن أن تكافح مع الفئات المعقدة، وأن جودة كتاب الرموز تؤثر بشكل كبير على أداء LLM. يهدف الإطار المقترح إلى توفير نهج منهجي للباحثين لتعزيز موثوقية وصلاحية التصنيفات المعتمدة على LLM في أبحاث العلوم الاجتماعية، مما يوجه في النهاية اختيار وضبط LLMs لمهام محددة.
DOI: https://doi.org/10.1017/pan.2025.10017
Publication Date: 2025-09-19
Author(s): Andrew Halterman et al.
Primary Topic: Artificial Intelligence in Law
Overview
This research paper section discusses the integration of generative large language models (LLMs) in the coding of political texts, emphasizing the need for empirical validation of their performance against established codebook operationalizations. The authors present a five-stage framework for evaluating LLMs in this context, which includes preparing codebooks, testing LLM capabilities, assessing zero-shot accuracy, analyzing errors, and conducting supervised training. They empirically demonstrate this framework using three curated political science codebooks related to protest events, political violence, and manifestos, revealing that while current open-weight LLMs struggle with zero-shot adherence to codebooks, supervised instruction-tuning can significantly enhance their performance.
In conclusion, the authors highlight the distinct challenges of employing LLMs for codebook-based measurement, which necessitates adherence to detailed coding criteria rather than relying solely on pretraining. They advocate for the use of their five-stage framework to guide researchers in evaluating LLMs for specific measurement tasks, noting that as LLM technology advances, measurement accuracy is expected to improve. The paper contributes valuable datasets and methodological guidance, enabling researchers to make informed decisions regarding codebook preparation, model selection, and the choice between zero-shot and supervised approaches.
Introduction
In recent years, political scientists have increasingly utilized generative large language models (LLMs) for text analysis, particularly for their ability to classify documents in a “zero-shot” manner without labeled examples. This method involves providing LLMs with categorical labels and brief descriptions to generate predictions for document classifications. However, this approach raises concerns regarding measurement validity in political science, as it assumes that simple instructions can accurately operationalize complex concepts. The authors highlight that LLMs, trained on extensive corpora, may exhibit reasonable predictive accuracy for broadly defined concepts but can misinterpret specific definitions, leading to potential inaccuracies in classification.
To address these issues, the authors propose a five-stage evaluation framework aimed at enhancing codebook-LLM measurement. This framework includes stages for codebook preparation, label-free testing, zero-shot evaluation, robustness assessments, and supervised training of LLMs. They introduce seven new codebook-specific behavioral tests that can be applied early in research projects without requiring ground-truth labels. Additionally, the authors present three real-world political science codebooks related to contentious politics, political violence, and party manifestos, which pose challenges for LLMs due to their complexity. Their empirical findings reveal limitations in LLMs’ adherence to codebook instructions in zero-shot scenarios, but they demonstrate that supervised instruction-tuning can significantly enhance performance, improving classification accuracy by up to 55%. The paper concludes with guidance for researchers on implementing these training techniques effectively.
Results
In this section, the performance of four large language models (LLMs)—Mistral-7B, Mistral-NeMo-12B, and Llama-8B—was evaluated using behavioral tests based on the BFRS codebook. The models demonstrated strong performance on Tests I through III, which assessed legal outputs and the recovery of definitions and examples. However, all models exhibited sensitivity to the order of category descriptions in the codebook, suggesting potential attention issues in processing long prompts. Zero-shot evaluations revealed low weighted F1 scores, with Mistral-7B and Llama-8B achieving scores of 0.21 on Manifestos and marginal scores of 0.65 and 0.57 on CCC and BFRS, respectively. This indicates that the LLMs struggle with complex codebook tasks, necessitating improvements in their ability to comprehend and follow instructions.
Further analysis through ablation experiments indicated that the systematic removal of components led to decreased performance, highlighting the importance of these components for achieving baseline F1 scores. Notably, omitting class definitions sometimes improved performance, suggesting that the models may not fully adhere to label definitions. Manual analysis revealed that while Mistral-7B showed high compliance with prompt instructions for BFRS and CCC, it exhibited significant non-compliance for Manifestos, including hallucinations and incorrect label predictions based on lexical overlap rather than accurate comprehension of codebook definitions. Instruction-tuning efforts yielded improvements in performance, particularly for Mistral-7B, which saw a 55% relative increase in F1 score on BFRS after fine-tuning. These findings underscore the potential of instruction-tuning to enhance LLM performance on challenging codebook tasks, albeit with increased computational demands.
Discussion
In this section, the authors present a comprehensive framework for evaluating codebook-LLM measurement, focusing on the multi-class classification of documents using large language models (LLMs) guided by a human-written codebook. The framework consists of five stages: (1) **Codebook Preparation**, where researchers operationalize social science concepts into a semi-structured format; (2) **Label-Free Behavioral Testing**, which assesses an LLM’s ability to follow instructions without requiring labeled data; (3) **Zero-Shot Evaluation with Labels**, where a small labeled dataset is used to evaluate the LLM’s performance; (4) **Zero-Shot Error Analysis**, which involves systematic error analysis and ablation studies to understand the LLM’s weaknesses; and (5) **Supervised Fine-Tuning**, where researchers may opt to fine-tune the LLM if zero-shot performance is insufficient.
The authors highlight that while previous studies have demonstrated LLMs’ potential in classifying social science concepts, their framework emphasizes the importance of precise operationalizations and structured codebooks to improve classification accuracy. They note that even high-performing LLMs can struggle with complex categories, and that the quality of the codebook significantly influences the LLM’s performance. The proposed framework aims to provide a systematic approach for researchers to enhance the reliability and validity of LLM-based classifications in social science research, ultimately guiding the selection and tuning of LLMs for specific tasks.
