تحويل معايير نجاح الوصول إلى الويب اليدوية إلى تلقائية: نهج قائم على LLM
Turning manual web accessibility success criteria into automatic: an LLM-based approach

المجلة: Universal Access in the Information Society، المجلد: 24، العدد: 1
DOI: https://doi.org/10.1007/s10209-024-01108-z
تاريخ النشر: 2024-03-16
المؤلف: Juan Miguel López وآخرون
الموضوع الرئيسي: الوصول الرقمي لذوي الإعاقة

نظرة عامة

تبحث الدراسة في إمكانيات أنظمة نماذج اللغة الكبيرة (LLM) لتقييم معايير نجاح الوصول إلى الويب التي تتطلب تقليديًا تقييمًا يدويًا. تركز الدراسة بشكل خاص على ثلاثة معايير من إرشادات الوصول إلى محتوى الويب (WCAG): 1.1.1 المحتوى غير النصي، 2.4.4 غرض الرابط (في السياق)، و3.1.2 لغة الأجزاء. طورت الدراسة نصوصًا تعتمد على LLM لتقييم هذه المعايير وقارنت أدائها مع أدوات تقييم الوصول الآلي الموجودة. أظهرت النتائج أنه بينما غالبًا ما تفشل الأدوات الآلية في اكتشاف المشكلات بشكل موثوق، حققت النصوص المعتمدة على LLM معدل اكتشاف إجمالي مثير للإعجاب بلغ 87.18%، حيث نجحت في تحديد مشكلات الوصول التي تم تجاهلها من قبل المقيمين التقليديين.

تشير النتائج إلى أن LLM يمكن أن تعزز بشكل كبير اختبار الوصول الآلي، مما يوفر تقييمًا أكثر شمولاً للامتثال لمعايير WCAG. تدعو الدراسة إلى مزيد من البحث لاستكشاف استخدام LLM مفتوحة المصدر، والتي يمكن أن تخفف من التكاليف ومخاوف الخصوصية المرتبطة بالنماذج الملكية. بالإضافة إلى ذلك، تؤكد على الحاجة إلى تطوير مطالبات عامة لتمكين LLM من معالجة مجموعة واسعة من اختبارات الوصول، مما يقلل من الاعتماد على الحلول المخصصة ويحسن الفعالية العامة لتقييمات الوصول إلى الويب. تضع هذه الدراسة الأساس لتطوير منهجيات تقييم الوصول الآلي، مما يعود بالنفع في النهاية على المطورين والمنظمات في إنشاء محتوى رقمي أكثر شمولاً.

مقدمة

تؤكد مقدمة الورقة على الأهمية الحاسمة لوصول الويب، والذي يهدف إلى القضاء على الحواجز التي تمنع الأفراد ذوي الإعاقات من استخدام المواقع الإلكترونية بشكل فعال. تسلط الضوء على مشكلات الوصول المستمرة، مثل البدائل النصية غير الكافية للصور وسوء تباين الألوان، والتي تعيق تفاعل المستخدم. توفر إرشادات الوصول إلى محتوى الويب (WCAG) إطارًا لإنشاء محتوى رقمي يمكن الوصول إليه، موضحة معايير النجاح القابلة للقياس (SC) عبر أربعة مبادئ: القابلية للإدراك، القابلية للتشغيل، القابلية للفهم، والصلابة. على الرغم من إصدار WCAG 2.2، فإن اختبار الامتثال يلتزم بشكل أساسي بـ WCAG 2.1 السابقة، مما يتطلب نهجًا مختلطًا للتقييمات الشاملة التي تجمع بين الأدوات الآلية، واختبار المستخدم، وتدقيق الخبراء لتحديد حواجز الوصول.

تستكشف الورقة أيضًا إمكانيات الذكاء الاصطناعي التوليدي (GenAI)، وخاصة نماذج اللغة الكبيرة (LLMs)، في أتمتة تقييم معايير نجاح WCAG التي تتطلب حاليًا تقييمًا يدويًا. يطرح المؤلفون سؤال بحث يتعلق بفعالية LLM في هذا السياق ويقدمون نتائج من دراسة محكومة. تشير النتائج إلى أنه بينما تظهر أدوات الاختبار الآلي التقليدية دقة محدودة (0% إلى 59%)، حققت النصوص المعتمدة على LLM التي تم تطويرها لهذه الدراسة معدل اكتشاف مثير للإعجاب بلغ 87.18%، حيث نجحت في تحديد المشكلات التي تم تجاهلها من قبل الأدوات الموجودة. يشير هذا إلى أن LLM يمكن أن تعزز بشكل كبير تقييمات الوصول إلى الويب، مما يقلل من الوقت والتكاليف المرتبطة بالتقييمات اليدوية. توضح الورقة هيكلها، موضحة الأعمال ذات الصلة، والمنهجية، والنتائج، والآثار على الأبحاث المستقبلية.

الطرق

توضح قسم الطرق نهج التقييم لتحديد قدرة نماذج اللغة الكبيرة (LLMs) على أتمتة اختبار الوصول وفقًا لمعايير نجاح إرشادات الوصول إلى محتوى الويب (WCAG) المحددة. يبدأ بتحديد السياق والدافع للدراسة، مشددًا على ضرورة الفحوصات اليدوية لبعض المعايير. يحدد المؤلفون المعايير المحددة من WCAG التي تم اختيارها كحالات اختبار ويشرحون المواد المستخدمة، والتي تشمل نموذج كود HTML، ومقيمات الوصول، وقواعد WCAG-ACT، وLLMs نفسها.

يصف القسم أيضًا الإطار الإجرائي المستخدم في الدراسة، موضحًا كيف تم تطوير نصوص تعتمد على LLM لكل معيار WCAG مختار ثم تقييمها مقابل حالات الاختبار. يتم إجراء هذا التقييم بالمقارنة مع المقيمين الآليين الموجودين، مما يمهد الطريق لتحليل شامل للنتائج في القسم التالي.

النتائج

يقدم قسم النتائج النتائج المستخلصة من تقييم قواعد ACT المختلفة، بما في ذلك حالات الاختبار المختارة والمعدلة. يتم تنظيم البيانات في سلسلة من الجداول (الجداول 3 إلى 8)، حيث يتوافق كل جدول مع معايير نجاح محددة (SC) من إرشادات الوصول إلى محتوى الويب (WCAG). يتضمن كل جدول أعمدة لرقم اختبار ACT، والنتائج المتوقعة، والتقييمات من أدوات وصول الويب المتعددة (A11y، Pa11y، Mauve++، AChecker، AccessMonitor، وLighthouse). يتم تصنيف النتائج على أنها ناجحة (P)، فاشلة (F)، لا يمكن تحديدها (CT)، غير موجودة (NP)، ولم يتم فحصها (NC)، مع أعمدة إضافية تشير إلى ما إذا كانت الأدوات قد تطابقت بدقة مع النتائج المتوقعة للحالات المختارة والمعدلة.

تتمثل الملاحظة الحرجة التي تم تسليط الضوء عليها في النتائج في حدوث إيجابيات زائفة، حيث يتم وضع علامات على الحالات المعدلة التي يجب أن تفشل بسبب مشكلات وصول متعمدة بشكل غير صحيح على أنها ناجحة. على سبيل المثال، قد تتلقى حالة معدلة تحتوي على نص بديل غير مناسب لصورة تقييم “ناجح”، مما يفشل في تمييزها عن الحالة المختارة المقابلة التي تلبي معايير الوصول. يوفر القسم أيضًا مقاييس مثل “النسبة المئوية المتوقعة الصحيحة” و”النسبة المئوية المتوقعة الدقيقة”، والتي تحدد أداء الأدوات في تحديد الحالات القابلة للوصول وغير القابلة للوصول بشكل متعمد، على التوالي. بشكل عام، تؤكد النتائج على قيود الأدوات التي تم تقييمها في تقييم الوصول إلى الويب بدقة، خاصة في التمييز بين الأمثلة الصحيحة وغير الصحيحة.

المناقشة

في هذا القسم، يناقش المؤلفون قيود أدوات تقييم الوصول إلى الويب الحالية في تقييم الامتثال لمعايير نجاح إرشادات الوصول إلى محتوى الويب (WCAG) المحددة، وخاصة تلك التي تتطلب حكمًا بشريًا. يبرز المؤلفون أنه بينما يمكن للأدوات الآلية تحديد العديد من مشكلات الوصول، فإنها غالبًا ما تفشل في اكتشاف المشكلات الدقيقة، مما يؤدي إلى إيجابيات وسلبيات زائفة. على سبيل المثال، قد تقوم الأدوات بشكل غير صحيح بوضع علامة على صفحة ويب على أنها متوافقة عندما لا يصف النص البديل لصورة المحتوى المرئي بدقة، أو عندما لا تتطابق أوصاف الروابط مع الوجهة. تؤكد هذه الفجوة على الحاجة إلى منهجية تقييم أكثر قوة تتضمن التحقق البشري جنبًا إلى جنب مع التقييمات الآلية.

يقترح المؤلفون الاستفادة من نماذج اللغة الكبيرة (LLMs) مثل ChatGPT وClaude وBard لتعزيز تقييم ثلاثة إرشادات WCAG محددة: 1.1.1 (المحتوى غير النصي)، 2.4.4 (غرض الرابط)، و3.1.2 (لغة الأجزاء). يوضحون نهج إثبات المفهوم حيث يتم استخدام LLMs لتحليل المحتوى ومقارنته، مما يوفر تقييمًا أكثر دقة للامتثال. على سبيل المثال، يمكن لـ LLMs توليد أوصاف للصور ومقارنتها بالنص البديل المقدم، أو تلخيص المحتوى المرتبط للتحقق من ملاءمة أوصاف الروابط. يجادل المؤلفون بأن هذا التكامل لـ LLMs يمكن أن يحسن بشكل كبير دقة تقييمات الوصول، مما يعزز في النهاية بيئة ويب أكثر شمولاً للمستخدمين ذوي الإعاقات.

القيود

تظهر الدراسة أن نماذج اللغة الكبيرة (LLMs) يمكن أن تعزز اختبار الوصول الآلي من خلال تحديد المشكلات التي غالبًا ما تتجاهلها اختبارات البرمجيات التقليدية. ومع ذلك، هناك عدة قيود تستدعي النظر في الأبحاث المستقبلية. أولاً، كان تقييم الدراسة مقصورًا على مجموعة اختبار WCAG ACT، مما يشير إلى الحاجة إلى اختبار أوسع عبر مجموعة أكبر من صفحات الويب الواقعية للتحقق من النتائج. ثانيًا، يشير التركيز على عدد محدود من معايير النجاح (SC) إلى أنه بينما تظهر LLMs وعدًا في تقييم الوصول إلى الويب، هناك حاجة لمزيد من التحقيق لتقييم قابليتها للتطبيق عبر المشهد المتنوع والمعقد لمعايير الوصول إلى الويب.

بالإضافة إلى ذلك، تسلط الدراسة الضوء على ضرورة تحديد أفضل LLMs والمطالبات لمهام محددة، حيث لم يتم اختبار جميع النماذج المحتملة. كانت التحليلات مقصورة على محتوى HTML، مما استبعد تنسيقات أخرى مثل PDFs أو النصوص الديناميكية التي تغطيها أيضًا إرشادات WCAG. أخيرًا، تشير الدراسة إلى أن القيم الحدية لتحديد نجاح الاختبار قد تحتاج إلى مزيد من التنقيح لتحقيق توازن مناسب بين الدقة والاسترجاع في تقييمات الوصول.

Journal: Universal Access in the Information Society, Volume: 24, Issue: 1
DOI: https://doi.org/10.1007/s10209-024-01108-z
Publication Date: 2024-03-16
Author(s): Juan Miguel López et al.
Primary Topic: Digital Accessibility for Disabilities

Overview

The research investigates the potential of large language model (LLM)-based systems to evaluate web accessibility success criteria that traditionally require manual assessment. Specifically, it focuses on three criteria from the Web Content Accessibility Guidelines (WCAG): 1.1.1 Non-text Content, 2.4.4 Link Purpose (In Context), and 3.1.2 Language of Parts. The study developed LLM-based scripts to evaluate these criteria and compared their performance against existing automated accessibility evaluators. The results indicated that while automated tools often failed to reliably detect issues, the LLM-based scripts achieved an impressive overall detection rate of 87.18%, successfully identifying accessibility problems that were overlooked by conventional evaluators.

The findings suggest that LLMs can significantly enhance automated accessibility testing, providing a more thorough evaluation of compliance with WCAG standards. The study advocates for further research to explore the use of open-source LLMs, which could mitigate costs and privacy concerns associated with proprietary models. Additionally, it emphasizes the need for developing generalized prompts to enable LLMs to address a wider array of accessibility tests, thereby reducing reliance on tailored solutions and improving the overall efficacy of web accessibility evaluations. This research lays the groundwork for advancing automated accessibility assessment methodologies, ultimately benefiting developers and organizations in creating more inclusive digital content.

Introduction

The introduction of the paper emphasizes the critical importance of web accessibility, which aims to eliminate barriers preventing individuals with disabilities from effectively using websites. It highlights persistent accessibility issues, such as inadequate text alternatives for images and poor color contrast, which hinder user interaction. The Web Content Accessibility Guidelines (WCAG) provide a framework for creating accessible digital content, outlining measurable success criteria (SC) across four principles: Perceivability, Operability, Understandability, and Robustness. Despite the release of WCAG 2.2, conformance testing primarily adheres to the earlier WCAG 2.1, necessitating a mixed-methods approach for comprehensive evaluations that combine automated tools, user testing, and expert audits to identify accessibility barriers.

The paper further explores the potential of Generative Artificial Intelligence (GenAI), particularly Large Language Models (LLMs), in automating the evaluation of WCAG success criteria that currently require manual assessment. The authors pose a research question regarding the efficacy of LLMs in this context and present findings from a controlled study. The results indicate that while traditional automated testing tools exhibit limited accuracy (0% to 59%), the LLM-based scripts developed for this study achieved an impressive 87.18% detection rate, successfully identifying issues overlooked by existing tools. This suggests that LLMs could significantly enhance web accessibility evaluations, reducing time and costs associated with manual assessments. The paper outlines its structure, detailing related work, methodology, results, and implications for future research.

Methods

The Methods section outlines the evaluation approach for determining the capability of Large Language Models (LLMs) to automate accessibility testing in accordance with specific Web Content Accessibility Guidelines (WCAG) success criteria. It begins by establishing the context and motivation for the study, highlighting the necessity of manual checks for certain criteria. The authors identify the specific WCAG criteria chosen as test cases and detail the materials utilized, which include sample HTML code, accessibility evaluators, WCAG-ACT rules, and the LLMs themselves.

The section further describes the procedural framework employed in the study, detailing how LLM-based scripts were developed for each selected WCAG criterion and subsequently evaluated against the test cases. This evaluation is conducted in comparison to existing automated evaluators, setting the stage for a comprehensive analysis of the results in the subsequent section.

Results

The results section presents the findings from the evaluation of various ACT rules, encompassing both selected and modified test cases. The data is organized in a series of tables (Tables 3 to 8), where each table corresponds to specific Success Criteria (SC) from the Web Content Accessibility Guidelines (WCAG). Each table includes columns for the ACT Test ID, expected results, and evaluations from multiple web accessibility tools (A11y, Pa11y, Mauve++, AChecker, AccessMonitor, and Lighthouse). The results are categorized as Passed (P), Failed (F), Cannot tell (CT), Not Present (NP), and Not Checked (NC), with additional columns indicating whether the tools accurately matched the expected outcomes for selected and modified cases.

A critical observation highlighted in the results is the occurrence of false positives, where modified cases that should fail due to intentional accessibility issues are incorrectly marked as passed. For instance, a modified case with inappropriate alternative text for an image may still receive a “Passed” evaluation, failing to distinguish it from the corresponding selected case that meets accessibility standards. The section also provides metrics such as “Valid Expected Percentage” and “Accurate Expected Percentage,” which quantify the tools’ performance in correctly identifying accessible and intentionally inaccessible cases, respectively. Overall, the findings underscore the limitations of the evaluated tools in accurately assessing web accessibility, particularly in distinguishing between valid and invalid examples.

Discussion

In this section, the authors discuss the limitations of current web accessibility evaluation tools in accurately assessing compliance with specific Web Content Accessibility Guidelines (WCAG) success criteria, particularly those that require human judgment. The authors highlight that while automated tools can identify many accessibility issues, they often fail to detect nuanced problems, leading to false positives and negatives. For instance, tools may incorrectly mark a webpage as compliant when the alternative text for an image does not accurately describe the visual content, or when link descriptions do not match the destination. This discrepancy underscores the need for a more robust evaluation methodology that incorporates human verification alongside automated assessments.

The authors propose leveraging Large Language Models (LLMs) like ChatGPT, Claude, and Bard to enhance the evaluation of three specific WCAG guidelines: 1.1.1 (Non-text Content), 2.4.4 (Link Purpose), and 3.1.2 (Language of Parts). They outline a proof-of-concept approach where LLMs are utilized to analyze and compare content, thereby providing a more accurate assessment of compliance. For example, LLMs can generate descriptions of images and compare them to the provided alt text, or summarize linked content to verify the appropriateness of link descriptions. The authors argue that this integration of LLMs could significantly improve the accuracy of accessibility evaluations, ultimately fostering a more inclusive web environment for users with disabilities.

Limitations

The research demonstrates that large language models (LLMs) can enhance automated accessibility testing by identifying issues that traditional software testing often overlooks. However, several limitations warrant consideration for future research. Firstly, the study’s evaluation was confined to the WCAG ACT Test suite, suggesting a need for broader testing across a more extensive array of real-world web pages to validate the findings. Secondly, the focus on a limited number of Success Criteria (SC) indicates that while LLMs show promise in web accessibility evaluation, further investigation is necessary to assess their applicability across the diverse and complex landscape of web accessibility standards.

Additionally, the research highlights the necessity of determining the optimal LLMs and prompts for specific tasks, as not all potential models were tested. The analysis was restricted to HTML content, leaving out other formats such as PDFs or dynamic scripts that are also covered by WCAG guidelines. Finally, the study notes that the threshold values for determining test success may need further refinement to achieve an appropriate balance between precision and recall in accessibility assessments.