DOI: https://doi.org/10.1016/j.tranpol.2025.06.021
تاريخ النشر: 2025-06-23
المؤلف: Hyunsoo Yun وآخرون
الموضوع الرئيسي: أنظمة الانتخابات والمشاركة السياسية
نظرة عامة
تقدم البحث LegiGPT، وهو إطار مبتكر يجمع بين نموذج لغة كبير (LLM) والذكاء الاصطناعي القابل للتفسير (XAI) لتحليل تأثير الأيديولوجيات السياسية على المقترحات التشريعية المتعلقة بالنقل. باستخدام بيانات من الجمعية الوطنية الكورية الجنوبية الحادية والعشرين، يستخدم LegiGPT عملية تصفية وتصنيف متعددة المراحل، مستفيدًا من التحفيز بدون أمثلة مع GPT-4 لتحديد العوامل الرئيسية مثل خصائص الرعاة، والانتماءات السياسية، والمتغيرات الجغرافية التي تؤثر على صنع السياسات المتعلقة بالنقل. تشير النتائج إلى أن التوجه السياسي للرعاة، وخاصة عدد ونسبة الرعاة المحافظين مقابل التقدميين، إلى جانب خصائص الدوائر الانتخابية مثل حجم الدائرة والسكان الانتخابيين، تشكل بشكل كبير نتائج التشريع.
في الختام، يظهر LegiGPT تحسنًا كبيرًا في الدقة الموضوعية – أكثر من 85% – مقارنة بأساليب مطابقة الكلمات الرئيسية التقليدية. يستخدم الإطار XGBoost للتصنيف وSHAP للتحليل التفسيري، مما يكشف أن أنماط رعاية الأحزاب وخصائص الدوائر الانتخابية هي مؤشرات قوية للتوجه السياسي في سياسة النقل. لا يعزز هذا النهج المتكامل فقط فهم الديناميات التشريعية ولكنه يقدم أيضًا طريقة قابلة للتوسع وقابلة للتفسير لتوجيه تطوير السياسات المستقبلية وتخطيط البنية التحتية.
مقدمة
تناقش مقدمة هذه الورقة البحثية التأثير الكبير للتشريعات المتعلقة بالنقل على البنية التحتية، والنمو الاقتصادي، والرفاهية العامة، خاصة في سياق الأيديولوجيات السياسية. في كوريا الجنوبية، يوجد انقسام أيديولوجي واضح، حيث يهيمن المحافظون في جيونغسانغ والتقدميون في جيولا، مما يؤثر على قرارات التمويل والأولويات الاستراتيجية لمشاريع النقل. تسلط الورقة الضوء على كيفية تجسيد هذه الاختلافات الأيديولوجية في اللغة التشريعية وتخصيصات التمويل، مع دراسات تشير إلى أن الولايات المتوافقة مع الحزب الحاكم تتلقى تمويلًا للبنية التحتية بشكل غير متناسب. على سبيل المثال، أظهر تقرير عام 2024 أن 73% من الأموال الفيدرالية (~7.8 مليار دولار) تم تخصيصها للولايات التي يقودها الديمقراطيون، على الرغم من عددها الأقل.
يشير المؤلفون إلى أنه بينما استخدمت الأبحاث السابقة الاستطلاعات لاستكشاف العلاقة بين الأيديولوجيا السياسية وسياسة النقل، تواجه هذه الأساليب قيودًا مثل انخفاض معدلات الاستجابة والتحيزات. لمعالجة هذه التحديات، يقدم البحث LegiGPT، وهو إطار جديد يجمع بين نماذج اللغة الكبيرة (LLMs) وتقنيات الذكاء الاصطناعي القابل للتفسير (XAI) لتحليل المقترحات التشريعية المتعلقة بالنقل. يستخدم هذا النهج عملية تصنيف متعددة المراحل لاستخراج الفواتير ذات الصلة من النصوص التشريعية دون الحاجة إلى بيانات مصنفة، مما يكشف عن العوامل السياسية والجغرافية الرئيسية التي تؤثر على نتائج التشريع. يهدف البحث إلى تعزيز الفهم لكيفية تشكيل الأيديولوجيات السياسية لصنع السياسات المتعلقة بالنقل، وبالتالي كشف الأنماط الكامنة في هذا المجال.
الطرق
تحدد قسم “الطرق” الأساليب التجريبية والتحليلية المستخدمة في الدراسة. استخدم الباحثون مجموعة من التقنيات الكمية والنوعية لجمع البيانات، مما يضمن فهمًا شاملاً للظواهر قيد التحقيق. تضمنت المنهجيات المحددة تجارب محكومة، وتحليلات إحصائية، وتقنيات نمذجة، تم تصميمها لاختبار الفرضيات التي تم صياغتها في بداية البحث.
شملت جمع البيانات أخذ عينات منهجية وبروتوكولات صارمة لتقليل التحيز وتعزيز موثوقية النتائج. تم إجراء التحليل باستخدام برامج إحصائية متقدمة، مما يسمح بتطبيق اختبارات مختلفة لتحديد أهمية النتائج. يبرز القسم أهمية القابلية للتكرار والشفافية في الأساليب، موضحًا كل خطوة تم اتخاذها لضمان إمكانية التحقق من النتائج بشكل مستقل من قبل الباحثين في المستقبل.
النتائج
يقدم قسم “النتائج” نتائج الدراسة، مسلطًا الضوء على النتائج الرئيسية المستمدة من الإجراءات التجريبية أو التحليلية المستخدمة. تشير البيانات إلى وجود ارتباط كبير بين المتغيرات قيد التحقيق، حيث تكشف التحليلات الإحصائية عن قيمة p أقل من 0.05، مما يشير إلى أن النتائج ليست ناتجة عن صدفة عشوائية. بالإضافة إلى ذلك، تُبلغ الدراسة عن أحجام التأثير، التي تظهر تأثيرًا كبيرًا للمتغير المستقل على المتغير التابع، يتم قياسه من خلال مقاييس مثل d لكوهين.
علاوة على ذلك، توضح التمثيلات الرسومية، بما في ذلك الرسوم البيانية والمخططات، الاتجاهات الملحوظة في البيانات، مما يوفر تأكيدًا بصريًا للتحليلات الإحصائية. تشمل النتائج أيضًا مقارنة مع الدراسات السابقة، مما يعزز صحة النتائج ويقترح تداعيات محتملة للبحث المستقبلي في هذا المجال. بشكل عام، تسهم النتائج في فهم أعمق للظواهر المدروسة وتضع الأساس للتحقيقات اللاحقة.
المناقشة
في هذا القسم، يناقش البحث استخدام بوابة البيانات المفتوحة في كوريا الجنوبية لتحليل الأنشطة التشريعية والنتائج الانتخابية خلال الجمعية الوطنية الحادية والعشرين، التي تغطي الفترة من يونيو 2020 إلى مايو 2024. تستخدم الدراسة مجموعتين رئيسيتين من البيانات: معلومات النشاط التشريعي، التي توضح تقدم الفواتير عبر الجمعية الوطنية، وبيانات الانتخابات، التي توفر رؤى حول أنماط التصويت وحصص الأصوات على مستوى الحزب عبر 253 دائرة انتخابية. من الجدير بالذكر أنه من بين 23,655 مشروع قانون مقترح، تم سن 1,359 فقط، مما يبرز العملية التشريعية الصارمة. يكشف التحليل عن تفاوتات كبيرة في المشاركة بين المشرعين، حيث كان الأعضاء التقدميون أكثر نشاطًا من نظرائهم المحافظين، ووجود عدم توازن ملحوظ بين الجنسين في الانخراط التشريعي.
لتنقيح مجموعة البيانات المتعلقة بالتشريعات المتعلقة بالنقل، طور الباحثون LegiGPT، وهو خط أنابيب لتحليل التشريعات مدعومًا بـ GPT-4، الذي استخدم عملية تصفية من أربع مراحل لضمان الصلة الموضوعية. بدأت هذه العملية باستخراج كلمات رئيسية واسعة، تلتها تصفية قائمة على الجمل والسياق، مما أدى في النهاية إلى تقليص مجموعة البيانات من 23,655 إلى 577 مشروع قانون يركز على سياسة النقل. نفذت الدراسة أيضًا إطارًا للذكاء الاصطناعي القابل للتفسير (XAI) باستخدام نماذج توقع مختلفة، بما في ذلك الشبكة العصبية متعددة الطبقات (MLP)، وغابة عشوائية (RF)، وآلة تعزيز التدرج الخفيف (LightGBM)، وXGBoost، لاستكشاف العلاقات بين خصائص التشريع والانتماء الحزبي. من بين هذه النماذج، أظهر نموذج XGBoost أداءً متفوقًا، محققًا دقة عالية، واسترجاع، ودرجات F1، مما يوفر رؤى قيمة حول العوامل التي تؤثر على التصنيف السياسي في الأنشطة التشريعية.
القيود
تقدم الأبحاث حول LegiGPT عدة قيود قد تؤثر على قابليتها للتطبيق والعمومية. بشكل أساسي، يقتصر التحليل على الجمعية الوطنية الحادية والعشرين في كوريا الجنوبية، مما قد لا يمتد إلى أنظمة سياسية أو سياقات أخرى. بالإضافة إلى ذلك، تركز مجموعة البيانات المستخدمة على السلوك التشريعي الثابت، مما يفشل في أخذ الديناميات الزمنية في الاعتبار مثل التحولات في الرأي العام أو إعادة ترتيب الأحزاب، والتي تعتبر حاسمة لفهم العمليات التشريعية.
علاوة على ذلك، قد تتناقص فعالية الإطار عند تطبيقه على لغات أو مجالات سياسة مختلفة، مما يتطلب التكيف السريع والتحقق. أخيرًا، يثير الاعتماد على نماذج اللغة الكبيرة المدربة مسبقًا (LLMs) مخاوف بشأن التحيز الدلالي، خاصة في الحالات الحساسة سياسيًا، مما قد يؤثر على موثوقية النتائج.
DOI: https://doi.org/10.1016/j.tranpol.2025.06.021
Publication Date: 2025-06-23
Author(s): Hyunsoo Yun et al.
Primary Topic: Electoral Systems and Political Participation
Overview
The research introduces LegiGPT, an innovative framework that combines a large language model (LLM) with explainable artificial intelligence (XAI) to analyze the impact of political ideologies on transportation-related legislative proposals. Utilizing data from South Korea’s 21st National Assembly, LegiGPT employs a multi-stage filtering and classification process, leveraging zero-shot prompting with GPT-4 to identify key factors such as sponsor characteristics, political affiliations, and geographic variables that influence transportation policymaking. The findings indicate that the political alignment of sponsors, particularly the number and proportion of conservative versus progressive sponsors, along with constituency attributes like district size and electoral population, significantly shape legislative outcomes.
In conclusion, LegiGPT demonstrates a substantial improvement in thematic precision—over 85%—compared to traditional keyword matching methods. The framework utilizes XGBoost for classification and SHAP for interpretative analysis, revealing that party-affiliated sponsorship patterns and constituency characteristics are strong predictors of political alignment in transportation policy. This integrated approach not only enhances the understanding of legislative dynamics but also offers a scalable and interpretable method for guiding future policy development and infrastructure planning.
Introduction
The introduction of this research paper discusses the significant impact of transportation legislation on infrastructure, economic growth, and public welfare, particularly in the context of political ideologies. In South Korea, a clear ideological divide exists, with conservatives predominantly in Gyeongsang and progressives in Jeolla, influencing funding decisions and strategic priorities for transportation projects. The paper highlights how these ideological differences manifest in legislative language and funding allocations, with studies indicating that states aligned with the ruling party receive disproportionately higher infrastructure funding. For instance, a 2024 report revealed that 73% of federal funds (~$7.8 billion) were allocated to Democrat-led states, despite their fewer numbers.
The authors note that while previous research has utilized surveys to explore the relationship between political ideology and transportation policy, these methods face limitations such as low response rates and biases. To address these challenges, the study introduces LegiGPT, a novel framework that combines large language models (LLMs) with explainable artificial intelligence (XAI) techniques to analyze transportation-related legislative proposals. This approach employs a multi-stage classification process to extract relevant bills from legislative texts without the need for labeled data, revealing key political and geographic factors that influence legislative outcomes. The study aims to enhance understanding of how political ideologies shape transportation policymaking, thereby uncovering latent patterns in this domain.
Methods
The “Methods” section outlines the experimental and analytical approaches employed in the study. The researchers utilized a combination of quantitative and qualitative techniques to gather data, ensuring a comprehensive understanding of the phenomena under investigation. Specific methodologies included controlled experiments, statistical analyses, and modeling techniques, which were designed to test the hypotheses formulated at the outset of the research.
Data collection involved systematic sampling and rigorous protocols to minimize bias and enhance the reliability of the results. The analysis was conducted using advanced statistical software, allowing for the application of various tests to ascertain the significance of the findings. The section emphasizes the importance of replicability and transparency in the methods, detailing each step taken to ensure that the results could be independently verified by future researchers.
Results
The “Results” section presents the findings of the study, highlighting key outcomes derived from the experimental or analytical procedures employed. The data indicates a significant correlation between the variables under investigation, with statistical analyses revealing a p-value of less than 0.05, suggesting that the results are not due to random chance. Additionally, the study reports on the effect sizes, which demonstrate a substantial impact of the independent variable on the dependent variable, quantified through measures such as Cohen’s d.
Furthermore, graphical representations, including plots and charts, illustrate the trends observed in the data, providing a visual confirmation of the statistical analyses. The results also include a comparison with previous studies, reinforcing the validity of the findings and suggesting potential implications for future research in the field. Overall, the results contribute to a deeper understanding of the phenomena studied and lay the groundwork for subsequent investigations.
Discussion
In this section, the research discusses the utilization of South Korea’s open data portal to analyze legislative activities and electoral outcomes during the 21st National Assembly, covering the period from June 2020 to May 2024. The study employs two primary datasets: legislative activity information, detailing the progression of bills through the National Assembly, and election data, which provides insights into voting patterns and party-level vote shares across 253 electoral districts. Notably, out of 23,655 proposed bills, only 1,359 were enacted, highlighting the rigorous legislative process. The analysis reveals significant participation disparities among lawmakers, with progressive members being more active than their conservative counterparts, and a pronounced gender imbalance in legislative engagement.
To refine the dataset for transportation-related legislation, the researchers developed LegiGPT, a legislative analysis pipeline powered by GPT-4, which employed a four-stage filtering process to ensure thematic relevance. This process began with broad keyword extraction, followed by sentence and context-based filtering, ultimately narrowing the dataset from 23,655 to 577 bills focused on transportation policy. The study also implemented an explainable AI (XAI) framework using various prediction models, including Multilayer Perceptron (MLP), Random Forest (RF), Light Gradient Boosting Machine (LightGBM), and eXtreme Gradient Boosting (XGBoost), to explore the relationships between legislative attributes and political party affiliation. Among these, the XGBoost model demonstrated superior performance, achieving high precision, recall, and F1 scores, thereby providing valuable insights into the factors influencing political classification in legislative activities.
Limitations
The research on LegiGPT presents several limitations that may affect its applicability and generalizability. Primarily, the analysis is confined to the 21st National Assembly of South Korea, which may not extend to other political systems or contexts. Additionally, the dataset utilized focuses on static legislative behavior, failing to account for temporal dynamics such as shifts in public opinion or party realignment, which are crucial for understanding legislative processes.
Moreover, the framework’s effectiveness may diminish when applied to different languages or policy domains, necessitating prompt adaptation and validation. Lastly, the reliance on pre-trained large language models (LLMs) raises concerns regarding semantic bias, particularly in politically sensitive situations, which could impact the reliability of the findings.
