DOI: https://doi.org/10.36680/j.itcon.2025.019
تاريخ النشر: 2025-04-01
المؤلف: May Shayboun وآخرون
الموضوع الرئيسي: أبحاث الصحة والسلامة المهنية
نظرة عامة
تستعرض ورقة البحث تطبيق التعلم الآلي (ML) في تحليل تقارير الحوادث داخل صناعة البناء، مع تسليط الضوء على الاهتمام المتزايد في استخدام التعلم الآلي ومعالجة اللغة الطبيعية (NLP) لتعزيز السلامة ومنع الحوادث المهنية. على الرغم من الفوائد المحتملة، فإن اعتماد تقنيات التعلم الآلي كان محدودًا بسبب التحديات المتعلقة بتوافر البيانات، وهيكلة البيانات يدويًا، والاستخدام غير الكافي لأساليب التعلم غير المراقب. تصنف المراجعة الأدبيات بناءً على معالجة البيانات المسبقة، والخوارزميات، والتنفيذ، كاشفة عن فجوات كبيرة في استخدام مجموعات البيانات شبه المهيكلة وغير المهيكلة، مما يعقد استخراج الميزات وفعالية نماذج التعلم الآلي بشكل عام.
تشير النتائج إلى أنه بينما يمكن لنماذج التعلم الآلي تصنيف أسباب الحوادث وأنواعها والتنبؤ بشدتها، هناك قضايا مستمرة تتعلق بحدود البيانات والحاجة إلى مجموعات بيانات متوازنة. تؤكد المراجعة على ضرورة وجود طرق تطوير موحدة توضح كيف يمكن أن تُعلم توصيات التعلم الآلي اتخاذ القرار في سياقات السلامة. كما تدعو إلى إنشاء موارد معالجة اللغة الطبيعية الخاصة بالمجال، مثل المعاجم والأنتولوجيات المخصصة، لتحسين استخراج المعلومات ذات الصلة من تقارير الحوادث. يجب أن تركز الأبحاث المستقبلية على تعزيز الصرامة المنهجية لتطبيقات التعلم الآلي في سلامة البناء، واستكشاف تقنيات معالجة اللغة الطبيعية المتقدمة، والتحقق من النماذج لضمان قابليتها للتطبيق العملي وفعاليتها في السيناريوهات الواقعية.
مقدمة
تسلط المقدمة الضوء على المخاطر الكبيرة المرتبطة بقطاع البناء العالمي، وخاصة الحوادث المهنية ذات الحدوث العالي. وقد ركزت الأدبيات الحديثة بشكل متزايد على تطبيق تقنيات التعلم الآلي (ML) لتحليل وثائق صناعة البناء، خاصة في سياق السلامة. يتم استخدام تقنيات مختلفة، بما في ذلك التعلم العميق (DL) والتعلم الآلي، لاستخراج رؤى من البيانات المتعلقة بالحوادث الموجودة في التقارير والسجلات. يتم التأكيد على استخدام معالجة اللغة الطبيعية (NLP) لمهام مثل التعرف على الكيانات وتصنيف النصوص كنهج واعد لتحليل أسباب الحوادث وتحديد مؤشرات السلامة.
على الرغم من الفوائد المحتملة للتعلم الآلي والذكاء الاصطناعي (AI) في تعزيز أداء السلامة، فإن الحواجز مثل التكلفة والعملية والموثوقية تعيق اعتمادها على نطاق واسع. كما تنشأ تحديات من جودة البيانات الضعيفة، والحاجة إلى وضع علامات يدوية في التعلم العميق، والطبيعة الخاصة بالمجال لنماذج التعلم العميق. تم انتقاد الأدبيات لعدم وجود شفافية منهجية كافية في معالجة البيانات وتدريب نماذج التعلم الآلي، مما يعقد تقييم تأثير التعلم الآلي على عمليات السلامة. تهدف هذه الورقة إلى معالجة هذه الفجوات من خلال إجراء مراجعة مركزة لحالة الفن في تطبيقات التعلم الآلي لتحليل تقارير الحوادث في البناء، مع فحص معالجة البيانات المسبقة، والخوارزميات، والأساليب، وتحديات التنفيذ. تم هيكلة المراجعة حول مواضيع رئيسية، بما في ذلك معالجة البيانات المسبقة، واختيار الخوارزمية، واختبار النماذج والتحقق منها، مع التركيز على كل من عمليات التحقق الداخلية والخارجية.
نقاش
ت outlines قسم النقاش في ورقة البحث المنهجية والنتائج من مراجعة الأدبيات التي تركزت على تطبيق تقنيات التعلم الآلي (ML)، والتعلم العميق (DL)، وتنقيب البيانات (DM) لتحليل تقارير الحوادث في صناعة البناء. استخدمت المراجعة إطارًا مركزيًا حول المفاهيم، حيث تم البحث بشكل متكرر في أربعة قواعد بيانات وGoogle Scholar باستخدام كلمات مفتاحية مستهدفة. أسفر عملية الاختيار عن 37 مقالة ذات صلة، مما يبرز أهمية معالجة البيانات المسبقة، واختيار الخوارزميات، واختبار الأداء في سياق تحليل الحوادث. حددت المراجعة زيادة كبيرة في الأبحاث حول التعلم العميق وتنقيب البيانات في البناء منذ عام 2017، مما يبرز الأهمية المتزايدة لهذه التقنيات في تحسين نتائج السلامة.
تشير النتائج الرئيسية إلى أن معالجة البيانات المسبقة ضرورية لتعزيز كفاءة وفعالية نماذج التعلم الآلي، حيث تكون مجموعات البيانات المهيكلة أسهل في الإدارة مقارنة بتلك غير المهيكلة. تم استخدام خوارزميات متنوعة، بما في ذلك الغابات العشوائية (RF)، وXGBoost، وآلات الدعم الناقل (SVM)، بشكل متكرر، حيث غالبًا ما تتفوق النماذج المعتمدة على الأشجار على غيرها. كما لاحظت المراجعة التحديات في التعامل مع الفئات غير المتوازنة داخل مجموعات البيانات، مما يمكن أن يؤثر على دقة التصنيف. بشكل عام، تؤكد الورقة على ضرورة تحسين معالجة البيانات واختيار الخوارزميات لتعزيز القدرات التنبؤية للتعلم الآلي في تحليل حوادث البناء.
DOI: https://doi.org/10.36680/j.itcon.2025.019
Publication Date: 2025-04-01
Author(s): May Shayboun et al.
Primary Topic: Occupational Health and Safety Research
Overview
The research paper reviews the application of machine learning (ML) in analyzing accident reports within the construction industry, highlighting a growing interest in utilizing ML and natural language processing (NLP) to enhance safety and prevent occupational accidents. Despite the potential benefits, the adoption of ML techniques has been limited due to challenges in data availability, manual data structuring, and the underutilization of unsupervised learning methods. The review categorizes the literature based on data pre-processing, algorithms, and implementation, revealing significant gaps in the use of semi-structured and unstructured datasets, which complicates feature extraction and the overall effectiveness of ML models.
The findings indicate that while ML models can classify accident causes, types, and predict severity, there are persistent issues related to data limitations and the need for balanced datasets. The review emphasizes the necessity for standardized development methods that clarify how ML recommendations can inform decision-making in safety contexts. It also calls for the creation of domain-specific NLP resources, such as lexicons and customized ontologies, to improve the extraction of relevant information from accident reports. Future research should focus on enhancing the methodological rigor of ML applications in construction safety, exploring advanced NLP techniques, and validating models to ensure their practical applicability and effectiveness in real-world scenarios.
Introduction
The introduction highlights the significant risks associated with the global construction sector, particularly the high incidence of occupational accidents. Recent literature has increasingly focused on the application of machine learning (ML) techniques to analyze construction industry documentation, especially in the context of safety. Various technologies, including deep learning (DL) and ML, are being utilized to extract insights from accident-related data found in reports and registries. The use of natural language processing (NLP) for tasks such as entity recognition and text classification is emphasized as a promising approach for analyzing accident causes and identifying safety indicators.
Despite the potential benefits of ML and artificial intelligence (AI) in enhancing safety performance, barriers such as cost, practicality, and reliability hinder their widespread adoption. Challenges also arise from poor data quality, the need for manual labeling in DL, and the domain-specific nature of DL models. The literature has been critiqued for insufficient methodological transparency in processing data and training ML models, complicating the assessment of ML’s impact on safety processes. This paper aims to address these gaps by conducting a focused review of the state of the art in ML applications for accident report analysis in construction, examining data preprocessing, algorithms, methods, and implementation challenges. The review is structured around key themes, including data preprocessing, algorithm selection, and model testing and validation, with an emphasis on both internal and external validation processes.
Discussion
The discussion section of the research paper outlines the methodology and findings from a literature review focused on the application of machine learning (ML), deep learning (DL), and data mining (DM) techniques to analyze accident reports in the construction industry. The review utilized a concept-centric framework, iteratively searching four databases and Google Scholar with targeted keywords. The selection process resulted in 37 relevant articles, emphasizing the importance of data pre-processing, algorithm selection, and performance testing in the context of accident analysis. The review identified a significant increase in research on DL and DM in construction since 2017, highlighting the growing relevance of these technologies in improving safety outcomes.
Key findings indicate that data pre-processing is crucial for enhancing the efficiency and effectiveness of ML models, with structured datasets being easier to manage compared to unstructured ones. Various algorithms, including random forest (RF), XGBoost, and support vector machines (SVM), were frequently employed, with tree-based models often outperforming others. The review also noted challenges in handling unbalanced classes within datasets, which can affect classification accuracy. Overall, the paper underscores the necessity of optimizing data handling and algorithm selection to advance the predictive capabilities of ML in construction accident analysis.
