DOI: https://doi.org/10.1140/epjds/s13688-024-00481-2
تاريخ النشر: 2024-06-07
المؤلف: Marco Bronzini وآخرون
الموضوع الرئيسي: تقارير المسؤولية الاجتماعية للشركات
نظرة عامة
في السنوات الأخيرة، فرضت الهيئات التنظيمية الكشف عن المعلومات غير المالية من الشركات المدرجة في البورصة، مدفوعة بزيادة اهتمام المستثمرين بقضايا البيئة والمجتمع والحوكمة (ESG). الطبيعة غير المنظمة والمتنوعة للتقارير المستدامة تقدم تحديات في دمج البيانات للحصول على رؤى حول المسؤولية الاجتماعية للشركات (CSR). تتناول هذه الدراسة هذه التحديات من خلال استخدام نماذج اللغة الكبيرة (LLMs) ونموذج الجيل المعزز بالاسترجاع (RAG) لاستخراج رؤى هيكلية مرتبطة بـ ESG من التقارير المستدامة. كشفت التحليلات أن معايير ESG تشمل أكثر من 500 موضوع، وغالبًا ما تمتد إلى ما هو أبعد من التصنيفات الحالية، وأبرزت أوجه التشابه في الكشف بين الشركات داخل نفس المنطقة أو القطاع. ومن الجدير بالذكر أن النتائج تشير إلى أن الإفصاحات المتعلقة بـ ESG تؤثر بشكل كبير على تصنيفات ESG للشركات أكثر من البيانات المالية التقليدية.
تستفيد المنهجية المقترحة من نماذج اللغة الكبيرة التوليدية لتقييم إفصاحات موضوعات ESG بشكل منهجي، مما يساهم في مجال استخراج المعلومات التلقائي من التقارير المستدامة. من خلال استخدام نموذج لغة مدرب مسبقًا واعتماد تمثيل الرسم البياني الثنائي، تمكنت الدراسة من استخراج ثلاثيات هيكلية دلاليًا تقدم رؤى أعمق حول إجراءات ESG للشركات. لا تتجاوز هذه الطريقة قيود تقنيات استخراج المعلومات المفتوحة التقليدية (OIE) فحسب، بل تسهل أيضًا التحليلات الإحصائية المعنوية وقابلية تفسير درجات ESG. تشمل اتجاهات البحث المستقبلية دمج مصادر بيانات إضافية، مثل الأخبار المتعلقة بـ ESG، لتقييم التناقضات في ادعاءات الشركات واستخدام تصنيف الأدوار الدلالية (SRL) لإثراء المعلومات الهيكلية بالتفاصيل السياقية.
مقدمة
تتناول مقدمة هذه الورقة البحثية التحديات العالمية الملحة المتعلقة بالصحة العامة، وتغير المناخ، وعدم المساواة الاجتماعية، والحاجة إلى التنوع والشمول، مع التأكيد على أهمية الحلول التعاونية. تسلط الضوء على دور مبادئ البيئة والمجتمع والحوكمة (ESG) في مراقبة استدامة الشركات والممارسات الأخلاقية، والتي يتم تقييمها من خلال مؤشرات نوعية وكمية متنوعة. لا تسهل هذه المؤشرات فقط تقييم مساهمات الشركات في الأهداف الاجتماعية، بل تتماشى أيضًا مع أهداف التنمية المستدامة للأمم المتحدة (SDGs) الموضحة في أجندة 2030 للتنمية المستدامة. الطلب المتزايد على الإفصاحات غير المالية، المدفوع بالتنظيمات مثل توجيه الاتحاد الأوروبي للإبلاغ غير المالي (NFRD) وتوجيه الإبلاغ عن الاستدامة المؤسسية (CSRD)، يبرز ضرورة الشفافية في أداء ESG للشركات.
تقترح الورقة منهجية قائمة على البيانات تستخدم تقنيات معالجة اللغة الطبيعية (NLP) المتطورة لاستخراج وتحليل المبادرات المتعلقة بـ ESG تلقائيًا من تقارير استدامة الشركات. من خلال استخدام نماذج اللغة الكبيرة (LLMs) لاستخراج المعلومات وتمثيلات قائمة على الرسم البياني لتحليل البيانات، تهدف الدراسة إلى توضيح العلاقة بين الإفصاحات المؤسسية وتصنيفات ESG. تشمل المنهجية استخدام إطار SHapley Additive exPlanations (SHAP) لقابلية التفسير، مما يسمح بتمثيل هيكلي للإجراءات المتعلقة بـ ESG كـ ثلاثيات. تكشف النتائج أن الإفصاح الشامل يؤثر إيجابيًا على درجات ESG، بينما يمكن أن يكون التقرير المحدود ضارًا. بالإضافة إلى ذلك، تشير التحليلات إلى أن العوامل الخارجية، مثل التأثيرات الجغرافية والقطاعية، تلعب دورًا كبيرًا في تشكيل الإفصاحات المتعلقة بـ ESG. تسهم هذه الدراسة في الأدبيات المتعلقة بالاستدامة والمسؤولية الاجتماعية للشركات (CSR) من خلال توفير خط أنابيب NLP متقدم لاستخراج الرؤى من تقارير الاستدامة والتحقق من الفرضيات الحالية من خلال نهج قائم على البيانات.
طرق
في قسم الطرق، يحدد المؤلفون الإطار الشامل المستخدم في بحثهم، بدءًا من تحديد مصادر البيانات في القسم 3.1. يوضحون العمليات المعنية في إعداد البيانات (القسم 3.2)، وتوليد الثلاثيات (القسم 3.3)، والبناء اللاحق للرسم البياني المعرفي (KG) في القسم 3.4. ينتهي القسم بمناقشة الطرق التحليلية وتقنيات التقييم المستخدمة لتقييم ومقارنة الثلاثيات الناتجة، كما هو موضح في القسم 3.5. يضمن هذا النهج المنظم فحصًا منهجيًا للبيانات وفعالية المنهجيات المطبقة.
نتائج
في قسم النتائج، يقدم المؤلفون تحليلًا شاملاً لإحصائيات الشبكة المستمدة من ثلاثة رسوم بيانية ثنائية، مع التركيز على مقاييس مستوى العقد (القسم 4.1). تستكشف تحليل التنوع اللاحق التباين في إفصاحات موضوعات البيئة والمجتمع والحوكمة (ESG) عبر شركات وقطاعات مختلفة (القسم 4.2). تستمر التحليلات مع فحص أوجه التشابه بين الشركات بناءً على الإجراءات المعلنة بشكل مشترك (القسم 4.3)، تليها دراسة ما إذا كانت هذه الأوجه تتماشى مع معلومات أخرى عن الشركات (القسم 4.4).
بالإضافة إلى ذلك، يقيم المؤلفون قابلية تفسير درجات ESG من خلال تحديد الجوانب الواقعية الأكثر تأثيرًا (القسم 4.5). للحصول على رؤى إضافية، يتم تقديم تحليلات نوعية للثلاثيات الناتجة ودراسة إلغاء تتعلق بتعليمات النموذج في الملحقين A وB، على التوالي. يسمح هذا النهج المنظم بفهم دقيق للإفصاحات المتعلقة بـ ESG وآثارها عبر سياقات مختلفة.
مناقشة
في قسم المناقشة، تستعرض الورقة المشهد الحالي لطرق توليد الرسم البياني المعرفي (KG)، مع التأكيد على الانتقال من خطوط أنابيب معالجة اللغة الطبيعية التقليدية إلى استخدام نماذج اللغة الكبيرة (LLMs). تسلط الضوء على قيود تقنيات اكتساب المعرفة التقليدية، التي غالبًا ما تؤدي إلى تراكم الأخطاء بسبب التعلم المنفصل لمهام معالجة اللغة الطبيعية. يدعو المؤلفون إلى خط أنابيب NLP من مرحلة واحدة يدمج استخراج الكيانات والعلاقات، لا سيما من خلال طرق استخراج المعلومات المفتوحة (OIE) التي تولد ثلاثيات الموضوع-الفعل-المفعول (SPO). أظهرت التطورات الأخيرة في OIE، التي تستفيد من نماذج LLM القائمة على المحولات مثل BERT، وعدًا في تعزيز استخراج البيانات الهيكلية، لا سيما في سياق المعلومات المتعلقة بـ ESG.
تناقش الورقة أيضًا تطبيق نماذج LLM في تحليل البيانات النصية المتعلقة بـ ESG، مشيرة إلى دراسات متنوعة استخدمت NLP لاستخراج رؤى من تقارير استدامة الشركات. تقارن منهجيتها مع الأساليب الحالية من خلال التركيز على التوليد المشترك للرسم البياني المعرفي باستخدام نماذج LLM التوليدية، مما يسمح باستخراج ثلاثيات غنية دلاليًا موجهة نحو ESG. يتناول هذا النهج أوجه القصور في طرق OIE التقليدية، التي قد تتجاهل المعلومات الخاصة بالمجال. يستكشف المؤلفون أيضًا تأثير الأطر التنظيمية الإقليمية على الإفصاحات المتعلقة بـ ESG للشركات، كاشفين عن تحيزات في التقارير بناءً على السياقات الجغرافية والصناعية. من خلال توفير رؤى هيكلية مستمدة من الإفصاحات غير المالية للشركات، تهدف الورقة إلى المساهمة في المناقشات المستمرة في التمويل المستدام وتعزيز فهم المبادرات المتعلقة بـ ESG عبر قطاعات ومناطق مختلفة.
القيود
تسلط قيود هذه الدراسة الضوء على عدة مجالات رئيسية قد تؤثر على قوة ووضوح النتائج. أولاً، التحليل يعتمد على عينة محدودة من الشركات عبر قطاعات ومناطق مختلفة، مما يشير إلى أن توسيع مجموعة البيانات لتشمل حوالي 1,000 شركة قد يعزز من صحة النتائج. تستخدم عملية إعداد البيانات محلل PDF لاستخراج النص من تقارير الاستدامة، مما قد يؤدي إلى جمل غير صحيحة نحويًا تعقد عملية الاستخراج الدلالي. على الرغم من أن دمج نماذج اللغة الكبيرة (LLMs) وتقنيات الجيل المعزز بالاسترجاع (RAG) يخفف بعض هذه المشكلات – مما يؤدي إلى تغطية جمل فعالة بنسبة 68.1% – لا يزال هناك جزء كبير من البيانات يتم التخلص منه بسبب نقص التماسك. يمكن أن تستفيد الأبحاث المستقبلية من تحسين طرق تحليل الوثائق أو آليات تصفية إضافية لتحسين جودة البيانات.
علاوة على ذلك، فإن قابلية تفسير درجات البيئة والمجتمع والحوكمة (ESG)، المستمدة باستخدام قيم SHAP، تمثل قيدًا آخر. بينما توفر SHAP رؤى حول تأثير المتنبئين على التنبؤات الفردية، قد تؤدي القابلية العالمية المشتقة من الإحصائيات المجمعة إلى إخفاء العلاقات الخاصة بالقطاعات بسبب تنوع مجموعة البيانات. يشير هذا إلى أن تحليل قابلية التفسير القائم على مجموعات فرعية أكثر دقة قد ينتج عنه رؤى أوضح حول العوامل المؤثرة على مجموعات الشركات المحددة. أخيرًا، قد يؤدي الاعتماد على درجات ESG من منصة Refinitiv إلى إدخال تباين، حيث تستخدم وكالات التصنيف منهجيات مختلفة، مما قد يؤدي إلى تقييمات متباينة. يمكن أن تستكشف الدراسات المستقبلية سمات ESG إضافية من وكالات مختلفة لإثراء التحليل وتوفير فهم أكثر شمولاً لممارسات الاستدامة المؤسسية.
DOI: https://doi.org/10.1140/epjds/s13688-024-00481-2
Publication Date: 2024-06-07
Author(s): Marco Bronzini et al.
Primary Topic: Corporate Social Responsibility Reporting
Overview
In recent years, regulatory bodies have mandated the disclosure of non-financial information from publicly listed companies, driven by heightened investor interest in Environmental, Social, and Governance (ESG) issues. The unstructured and diverse nature of sustainability reports presents challenges in data integration for Corporate Social Responsibility (CSR) insights. This study addresses these challenges by employing Large Language Models (LLMs) and the Retrieval-Augmented Generation (RAG) paradigm to extract structured ESG-related insights from sustainability reports. The analysis revealed that ESG criteria encompass over 500 topics, often extending beyond existing categorizations, and highlighted disclosure similarities among companies within the same region or sector. Notably, the findings indicate that ESG disclosures significantly influence companies’ ESG ratings more than traditional financial data.
The proposed methodology leverages generative LLMs to systematically evaluate ESG topic disclosures, contributing to the field of automatic information extraction from sustainability reports. By utilizing a pre-trained language model and adopting a bipartite graph representation, the study enables the extraction of semantically structured triples that provide deeper insights into companies’ ESG actions. This approach not only overcomes limitations of traditional Open Information Extraction (OIE) techniques but also facilitates meaningful statistical analyses and interpretability of ESG scores. Future research directions include integrating additional data sources, such as ESG-related news, to assess inconsistencies in corporate claims and employing Semantic Role Labelling (SRL) to enrich the structured information with contextual details.
Introduction
The introduction of this research paper addresses the pressing global challenges of public health, climate change, social inequalities, and the need for diversity and inclusiveness, emphasizing the importance of collaborative solutions. It highlights the role of Environmental, Social, and Governance (ESG) principles in monitoring corporate sustainability and ethical practices, which are assessed through various qualitative and quantitative indicators. These indicators not only facilitate the evaluation of corporate contributions to societal goals but also align with the United Nations’ Sustainable Development Goals (SDGs) outlined in the 2030 Agenda for Sustainable Development. The increasing demand for non-financial disclosures, driven by regulations such as the European Union’s Non-Financial Reporting Directive (NFRD) and the Corporate Sustainability Reporting Directive (CSRD), underscores the necessity for transparency in corporate ESG performance.
The paper proposes a data-driven methodology utilizing state-of-the-art Natural Language Processing (NLP) techniques to automatically extract and analyze ESG initiatives from companies’ sustainability reports. By employing Large Language Models (LLMs) for information extraction and graph-based representations for data analysis, the study aims to elucidate the relationship between corporate disclosures and ESG ratings. The methodology includes the use of the SHapley Additive exPlanations (SHAP) framework for interpretability, allowing for a structured representation of ESG-related actions as triples. The findings reveal that comprehensive disclosure positively influences ESG scores, while limited reporting can be detrimental. Additionally, the analysis indicates that external factors, such as geographical and sectoral influences, play a significant role in shaping ESG disclosures. This work contributes to the literature on sustainability and Corporate Social Responsibility (CSR) by providing an advanced NLP pipeline for extracting insights from sustainability reports and validating existing hypotheses through a data-driven approach.
Methods
In the Methods section, the authors outline the comprehensive framework employed in their research, beginning with the identification of data sources in Section 3.1. They detail the processes involved in data preparation (Section 3.2), the generation of triples (Section 3.3), and the subsequent construction of the knowledge graph (KG) in Section 3.4. The section concludes with a discussion of the analytical methods and evaluation techniques utilized to assess and compare the generated triples, as presented in Section 3.5. This structured approach ensures a systematic examination of the data and the effectiveness of the methodologies applied.
Results
In the Results section, the authors present a comprehensive analysis of network statistics derived from three bipartite graphs, focusing on node-level metrics (Sect. 4.1). A subsequent diversity analysis explores the variation in Environmental, Social, and Governance (ESG) topic disclosures across different companies and sectors (Sect. 4.2). The analysis continues with an examination of company similarities based on jointly disclosed ESG actions (Sect. 4.3), followed by an investigation into whether these similarities correlate with other company information (Sect. 4.4).
Additionally, the authors assess the interpretability of ESG scores by identifying the most influential factual aspects (Sect. 4.5). For further insights, qualitative analyses of the generated triples and an ablation study concerning the model’s instruction are provided in Appendices A and B, respectively. This structured approach allows for a nuanced understanding of ESG disclosures and their implications across various contexts.
Discussion
In the discussion section, the paper reviews the current landscape of Knowledge Graph (KG) generation methods, emphasizing the transition from traditional Natural Language Processing (NLP) pipelines to the utilization of Large Language Models (LLMs). It highlights the limitations of conventional knowledge acquisition techniques, which often lead to error accumulation due to disjoint learning of NLP tasks. The authors advocate for a one-stage NLP pipeline that integrates entity and relation extraction, particularly through Open Information Extraction (OIE) methods that generate subject-predicate-object (SPO) triples. Recent advancements in OIE, leveraging transformer-based LLMs like BERT, have shown promise in enhancing the extraction of structured data, particularly in the context of Environmental, Social, and Governance (ESG) information.
The paper further discusses the application of LLMs in analyzing ESG-related textual data, noting various studies that have employed NLP to extract insights from corporate sustainability reports. It contrasts its methodology with existing approaches by focusing on the joint generation of KGs using generative LLMs, which allows for the extraction of semantically rich ESG-oriented triples. This approach addresses the shortcomings of traditional OIE methods, which may overlook domain-specific information. The authors also explore the influence of regional regulatory frameworks on corporate ESG disclosures, revealing biases in reporting based on geographical and industry-specific contexts. By providing structured insights derived from companies’ non-financial disclosures, the paper aims to contribute to ongoing discussions in sustainable finance and enhance the understanding of corporate ESG initiatives across different sectors and regions.
Limitations
The limitations of this study highlight several key areas that could affect the robustness and interpretability of the findings. Firstly, the analysis is based on a limited sample of companies across various sectors and regions, suggesting that expanding the dataset to include around 1,000 companies could enhance the validity of the results. The data preparation process utilizes a PDF parser to extract text from sustainability reports, which may lead to syntactically flawed sentences that complicate semantic extraction. Although the integration of large language models (LLMs) and retrieval-augmented generation (RAG) techniques mitigates some of these issues—resulting in a 68.1% effective sentence coverage—there remains a significant portion of data that is discarded due to lack of coherence. Future research could benefit from improved document parsing methods or additional filtering mechanisms to enhance data quality.
Moreover, the interpretability of Environmental, Social, and Governance (ESG) scores, derived using SHAP values, presents another limitation. While SHAP provides insights into the influence of predictors on individual predictions, the global interpretability derived from aggregated statistics may obscure sector-specific relationships due to the diversity of the dataset. This suggests that a more nuanced, subset-based interpretability analysis could yield clearer insights into the factors affecting specific company clusters. Lastly, reliance on ESG scores from the Refinitiv platform may introduce variability, as different rating agencies employ distinct methodologies, potentially leading to divergent assessments. Future studies could explore additional ESG attributes from various agencies to enrich the analysis and provide a more comprehensive understanding of corporate sustainability practices.
