DOI: https://doi.org/10.32604/cmes.2025.075442
تاريخ النشر: 2026-01-01
المؤلف: Van-Viet Nguyen وآخرون
الموضوع الرئيسي: تقنيات هندسة البرمجيات المدفوعة بالنماذج
نظرة عامة
تقدم هذه الورقة البحثية إطارًا موحدًا للتوليد الآلي والتحقق والتقييم لرسوم UML باستخدام نماذج اللغة الكبيرة (LLMs). يدمج الهيكل تطوير المتطلبات، وتوليد UML، والتحقق متعدد الوسائط، مستخدمًا LLaMA-3.2-1B-Instruct لتوليد متطلبات تركز على المستخدم وDeepSeek-R1-Distill-Qwen-32B لتحويل هذه المتطلبات إلى كود PlantUML. أنتجت الدراسة مجموعة بيانات صناعية تضم 11,997 رسم UML عبر ست عائلات رئيسية من الرسوم، محققة معدل نجاح تجميع قدره 89.5%. أظهر أسلوب التقييم متعدد الوسائط توافقًا قويًا بين التقييمات الآلية وتقييمات الخبراء، كما يتضح من معامل ارتباط بيرسون r = 0.82 وFleiss’ Kappa 0.78، مما يؤكد فعالية نظام التقييم المقترح.
يعالج الإطار التحديات المتعلقة بتوليد رسوم UML من خلال دمج حلقة تغذية راجعة لتحسين المخرجات ذات الجودة المنخفضة ونظام تحقق متعدد المنظورات باستخدام نماذج الرؤية واللغة. تهدف الأعمال المستقبلية إلى توسيع الإطار ليشمل أنواعًا إضافية من رسوم UML وتعزيز قدرات التفكير الزمني، خاصة لرسوم التسلسل. تسلط الأبحاث الضوء على الإمكانية لدمج التوليد المعزز بالاسترجاع (RAG) لتحسين الأساس السياقي وتقترح إثراء خط أنابيب التقييم بمقاييس متعددة الوسائط للحصول على رؤى أعمق. بشكل عام، تؤسس هذه الدراسة أساسًا قويًا لبناء مجموعات بيانات UML آليًا والتحقق متعدد الوسائط، مما يمهد الطريق للتقدم في نمذجة البرمجيات المدفوعة بالذكاء الاصطناعي وأتمتة التصميم.
مقدمة
تسلط مقدمة هذه الورقة البحثية الضوء على التأثير التحويلي لهندسة Transformer على هندسة البرمجيات، مع التأكيد على التعقيد المتزايد وحجم أنظمة البرمجيات الحديثة. بينما كانت لغة النمذجة الموحدة (UML) تسهل تقليديًا التواصل بين أصحاب المصلحة في تصميم البرمجيات، فقد تضاءل أهميتها في السنوات الأخيرة بسبب تفضيل منهجية Agile للرسوم البيانية الخفيفة. مع انتشار نماذج التعلم العميق، أعيدت الحاجة إلى تمثيلات UML قوية، ومع ذلك، لا يزال إنشاء رسوم UML يدويًا يتطلب جهدًا كبيرًا وغير متسق. تحدد الورقة مشكلتين رئيسيتين في البحث: التوليد الآلي لمجموعات بيانات UML الشاملة التي تشمل الجوانب الهيكلية والسلوكية، والتحقق من صحة هذه المجموعات من حيث الدقة والفائدة، خاصة في ضمان الاتساق بين الرسوم.
لمعالجة هذه التحديات، يقترح المؤلفون إطارًا جديدًا يدمج خط أنابيب خوارزمي لتوليد نماذج UML، وطريقة لتحويل هذه النماذج إلى مجموعات بيانات قابلة للقراءة آليًا، وإطار تحقق متعدد الوسائط يستفيد من التفكير المدفوع بالذكاء الاصطناعي. يهدف هذا الإطار إلى إنتاج مجموعة متنوعة من رسوم UML مع الحفاظ على النزاهة الدلالية وتسهيل عمليات التحقق الآلي. كما تؤكد الورقة على ندرة مجموعات بيانات UML عالية الجودة وقيود الأساليب الحالية، مما يبرز الحاجة إلى حل منهجي وآلي يعزز قابلية التوسع والموثوقية في نمذجة UML في هندسة البرمجيات. تشمل المساهمات إصدار مجموعة بيانات UML صناعية كبيرة، والتي تعمل كمعيار للبحوث المستقبلية في النمذجة الآلية وهندسة البرمجيات المدفوعة بالذكاء الاصطناعي.
النتائج
توفر نتائج الدراسة تقييمًا مفصلًا لفعالية الإطار المقترح في توليد والتحقق من ثلاثة أنواع من رسوم UML: حالات الاستخدام، والفئات، والتسلسل. أُجري التحليل على مجموعة بيانات صناعية تضم حوالي 12,000 عينة، حيث يبرز التحليل جانبين رئيسيين: أداء خط أنابيب التوليد المزدوج LLM ونظام التحقق متعدد الوسائط. تشير النتائج إلى أن Qwen2.5-VL-3B يتفوق باستمرار على النماذج الأخرى عبر أنواع الرسوم المختلفة، حيث يتفوق بشكل خاص في رسوم الفئات مع تركيز عالٍ من التوقعات عند الدرجة 6، مما يشير إلى دقة قوية في التقاط علاقات الفئات. بالمقابل، يظهر LLaMA-3.2-11B-Vision-Instruct دقة جزئية، خاصة في الرسوم السلوكية، حيث غالبًا ما يفشل في تحقيق أعلى الدرجات.
تكشف رسوم UML السلوكية عن تباين كبير في أداء النماذج. بالنسبة لرسوم حالات الاستخدام، يظهر LLaMA-3.2-11B-Vision-Instruct تركيزًا قويًا عند الدرجة 5 ولكنه يفتقر إلى التنوع، بينما يظهر Qwen2.5-VL-3B توزيعًا أكثر توازنًا. تسلط رسوم التسلسل الضوء على قوة Qwen2.5-VL-3B في نمذجة التفاعلات الديناميكية، بينما يواجه Aya-Vision-8B صعوبة في الاتساق. تؤكد رسوم الحالة قدرات Qwen2.5-VL-3B، حيث تحقق درجات عالية، بينما يظهر LLaMA-3.2-11B-Vision-Instruct وAya-Vision-8B أداءً غير متسق ومنخفض على التوالي. بشكل عام، تؤكد الدراسة على عدم التماثل الهيكلي-السلوكي في نمذجة UML، مما يشير إلى أنه بينما تكون الرسوم الهيكلية أكثر قابلية للتعامل مع نماذج الذكاء الاصطناعي/التعلم الآلي، فإن الرسوم السلوكية تمثل تحديات كبيرة، مما يتطلب مجموعات بيانات أغنى وأطر تحقق محسنة لتحسين توليد UML الآلي.
المناقشة
تسلط قسم المناقشة في الورقة الضوء على الفجوات الكبيرة في البحث الحالي حول مجموعات بيانات UML، خاصة بالمقارنة مع المجالات الأكثر تطورًا مثل معالجة اللغة الطبيعية ورؤية الكمبيوتر. غالبًا ما تكون مستودعات UML الحالية محدودة في النطاق والحجم، حيث تركز بشكل أساسي على الرسوم الهيكلية وتفتقر إلى تغطية سلوكية شاملة. بينما تم بذل بعض الجهود لاستخدام نماذج اللغة الطبيعية لاستخراج رسوم UML، تتطلب هذه الأساليب عادةً توجيهات محددة ولا توفر إطارًا موحدًا للتوليد على نطاق واسع. يؤكد المؤلفون على الحاجة إلى مجموعات بيانات UML شاملة ومتعددة الوسائط يمكن أن تدعم تطبيقات الذكاء الاصطناعي/التعلم الآلي المتقدمة، على غرار مجموعات البيانات التحويلية التي تُرى في مجالات أخرى، مثل ImageNet في رؤية الكمبيوتر.
يكشف تقييم أنواع مختلفة من رسوم UML عن خصائص أداء مميزة بين نماذج الرؤية-اللغة المستخدمة (VLMs). بالنسبة للرسوم السلوكية، مثل رسوم حالات الاستخدام ورسوم التسلسل، تظهر النماذج نقاط قوة متفاوتة في التقاط علاقات الممثل-النظام وتدفقات التفاعل. بشكل ملحوظ، يظهر Qwen2.5-VL-3B أداءً قويًا عبر درجات متعددة، بينما يظهر LLaMA-3.2-11B-Vision-Instruct قيودًا في تحقيق دقة التفاعل الكاملة. بالمقابل، يتم التعامل مع الرسوم الهيكلية، وخاصة رسوم الفئات، بشكل أكثر فعالية، حيث يحقق Qwen2.5-VL-3B دقة عالية في التقاط العلاقات الهيكلية. تؤكد النتائج على ضرورة تركيز الأبحاث المستقبلية على تعزيز دقة التعليق السلوكي وقدرات التفكير متعدد الوسائط، حيث تظل هذه الجوانب تحديات حاسمة في توليد UML الآلي.
القيود
تسلط قسم القيود الضوء على عدة قيود حاسمة في نتائج الدراسة. أولاً، يقتصر التقييم على مجموعة محدودة من أنواع رسوم UML—تحديدًا رسوم الفئات، والمكونات، والأشياء، وحالات الاستخدام، والتسلسل، والرسوم الحالة—مما يستبعد عائلات الرسوم المهمة الأخرى مثل رسوم النشاط، والنشر، والتواصل. قد يعيق هذا التركيز الضيق عمومية النتائج عبر المواصفات الكاملة لـ UML. ثانيًا، بينما تعتبر مجموعة البيانات المستخدمة متنوعة، إلا أنها صغيرة نسبيًا وتفتقر إلى عمق التعليق الموجود في المعايير المعتمدة في مجالات مثل رؤية الكمبيوتر ومعالجة اللغة الطبيعية. تكون هذه القيود واضحة بشكل خاص في مهام التفكير السلوكي التي تتضمن رسوم التسلسل والحالة، حيث تظهر جميع النماذج المختبرة نقاط ضعف.
بالإضافة إلى ذلك، يعتمد التقييم على ثلاثة هياكل تمثيلية فقط (Qwen2.5-VL-3B، LLaMA-3.2-11B-Vision-Instruct، وAya-Vision-8B)، والتي قد لا تمثل بشكل كافٍ الطيف الأوسع من نماذج الرؤية-اللغة والنماذج الموجهة نحو الكود. تعترف الدراسة بأن الأداء يمكن أن يكون حساسًا لتصميم المطالبات، مما يشير إلى أن التحقيق المنهجي في استراتيجيات المطالبات لرسوم UML لا يزال مسعى ضروريًا في المستقبل. علاوة على ذلك، بينما يتم استخدام درجات BLEU لتقييم الدقة النحوية، لم يتم استكشاف مقاييس هيكلية أخرى مثل مسافة تعديل الرسم أو التحقق القائم على التنفيذ بشكل كامل، مما قد يحد من اكتشاف التباينات الطوبوغرافية الدقيقة. كما تعطي الدراسة الأولوية للدقة والموثوقية على الأداء الحاسوبي، مما يشير إلى الحاجة إلى أعمال مستقبلية لتحليل مقاييس مثل زمن الاستدلال واستخدام الموارد. أخيرًا، بينما يعالج إطار التقييم الدقة النحوية والهيكلية، فإنه يقدم رؤى محدودة حول الدقة الدلالية، وهو أمر ضروري لضمان أن تمثيلات UML قابلة للتنفيذ وصحيحة وظيفيًا.
DOI: https://doi.org/10.32604/cmes.2025.075442
Publication Date: 2026-01-01
Author(s): Van-Viet Nguyen et al.
Primary Topic: Model-Driven Software Engineering Techniques
Overview
This research paper presents a unified framework for the automated generation, validation, and evaluation of Unified Modeling Language (UML) diagrams using large language models (LLMs). The architecture integrates requirement development, UML synthesis, and multimodal validation, employing LLaMA-3.2-1B-Instruct for generating user-focused requirements and DeepSeek-R1-Distill-Qwen-32B for converting these requirements into PlantUML code. The study produced a synthetic dataset of 11,997 UML diagrams across six major diagram families, achieving a compilation success rate of 89.5%. A multimodal scoring method demonstrated strong alignment between automated and expert evaluations, indicated by a Pearson correlation of r = 0.82 and a Fleiss’ Kappa of 0.78, confirming the effectiveness of the proposed scoring system.
The framework addresses the challenges of UML diagram generation by incorporating a feedback loop for refining low-quality outputs and a multi-perspective validation system using vision-language models. Future work aims to expand the framework to include additional UML diagram types and enhance temporal reasoning capabilities, particularly for Sequence Diagrams. The research highlights the potential for integrating retrieval-augmented generation (RAG) to improve contextual grounding and suggests enriching the evaluation pipeline with multimodal metrics for deeper insights. Overall, this study establishes a robust foundation for automated UML dataset construction and multimodal verification, paving the way for advancements in AI-driven software modeling and design automation.
Introduction
The introduction of this research paper highlights the transformative impact of the Transformer architecture on software engineering, emphasizing the increasing complexity and scale of modern software systems. While the Unified Modeling Language (UML) has traditionally facilitated communication among stakeholders in software design, its relevance has diminished in recent years due to the Agile methodology’s preference for lightweight diagrams. As deep learning models proliferate, the need for robust UML representations has re-emerged, yet the manual creation of UML diagrams remains labor-intensive and inconsistent. The paper identifies two core research problems: the automated generation of comprehensive UML datasets that encompass both structural and behavioral aspects, and the validation of these datasets’ correctness and utility, particularly in ensuring inter-diagram consistency.
To address these challenges, the authors propose a novel framework that integrates an algorithmic pipeline for generating UML models, a method for converting these models into machine-readable datasets, and a multimodal validation framework leveraging AI-driven reasoning. This framework aims to produce a diverse array of UML diagrams while maintaining semantic integrity and facilitating automated validation processes. The paper also emphasizes the scarcity of high-quality UML datasets and the limitations of existing approaches, underscoring the need for a systematic, automated solution that enhances the scalability and reliability of UML modeling in software engineering. The contributions include the release of a large-scale synthetic UML dataset, which serves as a benchmark for future research in automated modeling and AI-driven software engineering.
Results
The results of the study provide a detailed evaluation of the proposed framework’s effectiveness in generating and validating three types of UML diagrams: Use Case, Class, and Sequence. Conducted on a synthetic dataset of approximately 12,000 samples, the analysis highlights two main aspects: the performance of the dual-LLM generation pipeline and the multimodal validation system. The findings indicate that Qwen2.5-VL-3B consistently outperforms other models across various diagram types, particularly excelling in Class diagrams with a high concentration of predictions at score 6, indicating strong fidelity in capturing class relationships. In contrast, LLaMA-3.2-11B-Vision-Instruct demonstrates partial correctness, particularly in behavioral diagrams, where it often fails to achieve the highest scores.
Behavioral UML diagrams reveal significant variability in model performance. For Use Case diagrams, LLaMA-3.2-11B-Vision-Instruct shows a strong concentration at score 5 but lacks diversity, while Qwen2.5-VL-3B exhibits a more balanced distribution. Sequence diagrams highlight Qwen2.5-VL-3B’s robustness in modeling dynamic interactions, whereas Aya-Vision-8B struggles with coherence. State diagrams further emphasize Qwen2.5-VL-3B’s capabilities, achieving high scores, while LLaMA-3.2-11B-Vision-Instruct and Aya-Vision-8B show inconsistent and low performance, respectively. Overall, the study underscores a structural-behavioral asymmetry in UML modeling, indicating that while structural diagrams are more tractable for AI/ML models, behavioral diagrams present significant challenges, necessitating richer datasets and enhanced validation frameworks to improve automated UML synthesis.
Discussion
The discussion section of the paper highlights significant gaps in the current research on UML datasets, particularly in comparison to more developed fields like natural language processing and computer vision. Existing UML repositories are often limited in scope and scale, primarily focusing on structural diagrams and lacking comprehensive behavioral coverage. While some efforts have been made to utilize natural language models for UML diagram extraction, these approaches typically require specific orientations and do not provide a unified framework for large-scale generation. The authors emphasize the need for comprehensive, multimodal UML datasets that can support advanced AI/ML applications, akin to the transformative datasets seen in other domains, such as ImageNet in computer vision.
The evaluation of various UML diagram types reveals distinct performance characteristics among the employed Vision-Language Models (VLMs). For behavioral diagrams, such as use case and sequence diagrams, the models exhibit varying strengths in capturing actor-system relationships and interaction flows. Notably, Qwen2.5-VL-3B demonstrates robust performance across multiple scores, while LLaMA-3.2-11B-Vision-Instruct shows limitations in achieving complete interaction fidelity. In contrast, structural diagrams, particularly class diagrams, are handled more effectively, with Qwen2.5-VL-3B achieving high precision in capturing structural relationships. The findings underscore the necessity for future research to focus on enhancing behavioral annotation granularity and multimodal reasoning capabilities, as these aspects remain critical challenges in automated UML synthesis.
Limitations
The section on limitations highlights several critical constraints of the study’s findings. Firstly, the evaluation is confined to a limited selection of UML diagram types—specifically class, component, object, use case, sequence, and state diagrams—thus excluding other significant diagram families like activity, deployment, and communication diagrams. This narrow focus may hinder the generalizability of the results across the complete UML specification. Secondly, while the dataset utilized is diverse, it is relatively small and lacks the depth of annotation found in established benchmarks within fields such as computer vision and natural language processing. This limitation is particularly evident in behavioral reasoning tasks involving sequence and state diagrams, where all tested models exhibit weaknesses.
Additionally, the evaluation is based on only three representative architectures (Qwen2.5-VL-3B, LLaMA-3.2-11B-Vision-Instruct, and Aya-Vision-8B), which may not adequately represent the broader spectrum of vision-language and code-oriented models. The study acknowledges that performance can be sensitive to prompt design, suggesting that a systematic investigation into prompting strategies for UML diagrams remains a necessary future endeavor. Furthermore, while BLEU scores are employed to assess syntactic fidelity, other structural metrics such as graph edit distance or execution-based validation were not fully explored, potentially limiting the detection of nuanced topological discrepancies. The study also prioritizes accuracy and reliability over computational performance, indicating a need for future work to analyze metrics like inference time and resource utilization. Lastly, while the scoring framework addresses syntactic and structural fidelity, it offers limited insight into semantic correctness, which is essential for ensuring that UML representations are executable and functionally valid.
