schema-miner pro: ذكاء اصطناعي وكيل لتأسيس الأنطولوجيا عبر المخططات العلمية المكتشفة بواسطة LLM في سير عمل إنسان في الحلقة
schema-miner pro: Agentic AI for Ontology Grounding Over LLM-Discovered Scientific Schemas in a Human-in-the-Loop Workflow

شارك:
المجلة: Semantic Web، المجلد: 17، العدد: 3
DOI: https://doi.org/10.1177/22104968261431521
تاريخ النشر: 2026-04-25
المؤلف: Sameer Sadruddin وآخرون
الموضوع الرئيسي: الحوسبة العلمية وإدارة البيانات

نظرة عامة

القسم يقدم SCHEMA-MINER pro، وهو إطار عمل يعتمد على الإنسان في الحلقة مصمم لاستخراج وتأسيس المخططات الهيكلية من الأدبيات العلمية، مما يعالج تحدي تمثيل العمليات العلمية في صيغة قابلة للإدارة حسابيًا. من خلال دمج نماذج اللغة الكبيرة لاستخراج المخططات مع نظام قائم على الوكلاء، يتماشى الإطار مع العناصر المستخرجة إلى الأنطولوجيات الخارجية من خلال تفكير قابل للتفسير ومتعدد الخطوات. يستخدم الوكيل خوارزميات لغوية، تشابه دلالي، وتعليقات الخبراء لضمان دقة عملية التأسيس. تم عرض الإطار على سير العمل في تصنيع أشباه الموصلات، وبالتحديد ترسيب الطبقة الذرية (ALD) وحفر الطبقة الذرية (ALE)، بنجاح في ربط معلمات العملية والمخرجات بأنطولوجيا QUDT (الكميات، الوحدات، الأبعاد، والأنواع).

في الختام، يمثل SCHEMA-MINER pro تقدمًا كبيرًا في اكتشاف المخططات العلمية وتأسيس الأنطولوجيا، مع آثار تمتد إلى ما هو أبعد من تصنيع أشباه الموصلات إلى الويب الدلالي الأوسع. المخططات المستندة إلى الأنطولوجيا التي تم إنتاجها قابلة للتنفيذ آليًا وقابلة لإعادة الاستخدام عبر مجالات علمية متنوعة، مما يسهل سير العمل القابل للتكرار، والتفكير الآلي، والتشغيل البيني الدلالي. على سبيل المثال، يمكن أن تعزز مخطط ALD المستند إلى الأنطولوجيا من توضيح مجموعات البيانات التجريبية ودعم المقارنات المتسقة من حيث الوحدات. تشمل الاتجاهات المستقبلية مواءمة المخططات مع ويكيداتا لربط المعرفة على مستوى العملية بالكيانات العلمية الأوسع وتوسيع تطبيق الإطار ليشمل علم الأحياء والعلوم البيئية، مما يعالج قضايا مماثلة تتعلق بالتقارير غير المنظمة والتجزئة الدلالية. يهدف دمج SCHEMA-MINER pro مع مبادرات بيانات مفتوحة مرتبطة نشطة (LOD) والأشياء الرقمية FAIR إلى المساهمة في بنية تحتية منظمة وقابلة للاستعلام للعلوم المفتوحة.

مقدمة

تناقش مقدمة ورقة البحث الحاجة الملحة لاستخراج المعلومات الهيكلية من الأدبيات العلمية غير المنظمة لتعزيز نمذجة العمليات المعقدة في العالم الحقيقي. على الرغم من التقدم في طرق استخراج المعلومات (IE)، لا تزال الأساليب التقليدية تتطلب جهدًا كبيرًا وغالبًا ما تحتاج إلى مجموعات بيانات موسومة واسعة، خاصة في مجالات متخصصة مثل علوم المواد وعلوم الحياة. أظهرت التطورات الأخيرة في نماذج اللغة الكبيرة (LLMs) وعدًا في أتمتة استخراج البيانات الهيكلية، ومع ذلك لا تزال التحديات قائمة بسبب اللغة المتخصصة والمصطلحات المتنوعة عبر المنشورات. لمعالجة هذه القضايا، تقدم الورقة SCHEMA-MINER pro، وهو إطار عمل محسّن يدمج سير العمل القائم على الوكلاء مع LLMs لتأسيس الأنطولوجيا، مما يحسن من دقة وقابلية تفسير المخططات المستخرجة.

يستخدم SCHEMA-MINER pro نهجًا منهجيًا وتكراريًا يجمع بين مطابقة السلاسل الخوارزمية والبحث عن المتجهات الدلالية لمواءمة عناصر المخطط مع مفاهيم الأنطولوجيا، مع دمج التحقق من الخبراء طوال العملية. يتناقض هذا مع طرق التحفيز التقليدية الأحادية، التي غالبًا ما تفتقر إلى الشفافية وتتطلب تحققًا بشريًا واسعًا. يتم تقييم الإطار عبر عمليتين لأشباه الموصلات، ترسيب الطبقة الذرية (ALD) وحفر الطبقة الذرية (ALE)، مما يوضح فعاليته في التعامل مع المصطلحات العلمية المعقدة. بالإضافة إلى ذلك، تم تقديم واجهة ويب سهلة الاستخدام لتسهيل مشاركة الخبراء وتعزيز إمكانية الوصول إلى عملية اكتشاف المخططات. تؤكد النتائج على أهمية مشاركة الخبراء في تحسين المخططات، مشيرة إلى أن نهج التغذية الراجعة الهجين يحسن بشكل كبير من جودة وتماسك المعلومات المستخرجة.

طرق

استخدم البحث أداة تسمى SCHEMA-MINER، تم تطويرها بلغة بايثون باستخدام إطار عمل LangChain، لتسهيل استخراج المخططات من نماذج اللغة الكبيرة المختلفة (LLMs)، بما في ذلك GPT-4o وGPT-4turbo وLLaMA 3.1 (8B). تم إجراء التجارب على جهاز مزود بمعالج 16 نواة و32 جيجابايت من ذاكرة الوصول العشوائي، مع الاستفادة من خدمات السحابة لاستنتاجات LLM. قيمت الدراسة اكتشاف المخططات من خلال أربعة تكوينات تجريبية تتضمن أنواعًا مختلفة من تغذية الخبراء الراجعة – نصوص وصفية وتعديلات مباشرة على المخطط – عبر 21 تجربة. من الجدير بالذكر أن التكوين الذي جمع بين كلا نوعي التغذية الراجعة في كل تكرار أسفر عن أدق النتائج لكل من عمليات ترسيب الطبقة الذرية (ALD) وحفر الطبقة الذرية (ALE).

تم تسليط الضوء على أهمية أنطولوجيا الكميات والوحدات والأبعاد والأنواع (QUDT) في علوم المواد، مع التأكيد على دورها في توحيد التمثيلات القابلة للقراءة آليًا للكميات الفيزيائية. هذه التوحيد ضروري لضمان التناسق عبر مصادر البيانات المتنوعة، مما يمكّن من تحويل الوحدات بشكل آلي، ويعزز من قابلية تكرار البيانات التجريبية والتشغيل البيني. أشارت النتائج إلى أن GPT-4turbo تفوق على LLaMA 3.1 (8B) من حيث التماسك الهيكلي، كما يتضح من درجات ROUGE-L، مع الحفاظ على توافق دلالي قوي عبر النماذج، كما يتضح من درجات BERT العالية. بشكل عام، أظهر كل من GPT-4o وGPT-4turbo أداءً متفوقًا في التقاط تفاصيل مخططات ALD وALE التجريبية، مما يبرز قوتها في معالجة البيانات العلمية المعقدة.

نتائج

في هذا القسم، يتم تقييم أداء إطار عمل SCHEMA-MINER pro من خلال تحليلات كمية ونوعية، مع التركيز على عمليات ALD وALE. تقيم التقييم الكمي التباين والاختلافات الهيكلية في المخططات التي تم إنشاؤها بواسطة ثلاثة نماذج لغة كبيرة (LLMs): GPT-4o وGPT-4-turbo وLLaMA 3.1 (8B) عبر ثلاث مراحل من تحسين المخطط. يستخدم التقييم ثلاثة مقاييس – ROUGE-L وBLEU Score وBERTScore – لقياس تشابه المخطط، مع نتائج تشير إلى درجات متفاوتة من التوافق والتطور بين النماذج. جميع المخططات التي تم إنشاؤها متاحة للجمهور في المستودع.

يؤكد التقييم النوعي على استقرار مخرجات LLM، وهو أمر حاسم للحفاظ على تماسك المخطط أثناء التحسين. كشفت تعليقات الخبراء في المجال أن كل من GPT-4o وLLaMA 3.1 (8B) أظهرا استقرارًا عاليًا في عملية ALD، بينما أظهر GPT-4-turbo عدم استقرار من خلال إدخال خصائص غير موثقة. في عملية ALE، حافظ كل من GPT-4o وGPT-4-turbo على استقرار نسبي، بينما أنتجت LLaMA 3.1 (8B) العديد من الخصائص غير ذات الصلة، خاصة في المرحلة النهائية، مما أثر سلبًا على الدقة الدلالية للمخططات. تؤكد هذه النتائج على أهمية اختيار النموذج في تحقيق استخراج موثوق للمخططات.

مناقشة

تتناول قسم المناقشة في ورقة البحث التقدم والمساهمات التي يقدمها نظام SCHEMA-MINER pro، الذي يعزز استنتاج المخططات وتأسيس الأنطولوجيا في الأدبيات العلمية. بناءً على المنهجيات السابقة، يدمج SCHEMA-MINER pro نماذج اللغة الكبيرة (LLMs) مع سير العمل القائم على الإنسان في الحلقة لمعالجة تعقيدات النصوص العلمية، التي غالبًا ما تحتوي على مصطلحات كثيفة وعمليات معقدة. تستعرض الورقة الأعمال السابقة في اكتشاف المخططات، مشيرة إلى التطور من التقنيات القائمة على القواعد إلى الأساليب العصبية التي تستفيد من النماذج التوليدية لاستنتاج المخططات. من الجدير بالذكر أنها تؤكد على قيود الأنظمة الحالية، التي تركز عادةً على المجالات العامة وتفتقر إلى القدرة اللازمة للتكيف مع السياقات العلمية.

يميز نظام SCHEMA-MINER pro نفسه من خلال تقديم بنية معيارية تدعم سير العمل القائم على الذكاء الاصطناعي الوكلي، مما يسمح بإدارة ديناميكية لعملية استنتاج المخطط. يشمل ذلك دمج تغذية الخبراء الراجعة طوال سير العمل، مما يعزز من دقة وقابلية إعادة استخدام المخططات المنتجة. تناقش الورقة أيضًا دمج تأسيس الأنطولوجيا، حيث يتم مواءمة مكونات المخطط مع الأنطولوجيات المعتمدة لضمان التشغيل البيني وتسهيل تكامل المعرفة. النهج القائم على الوكلاء المستخدم في SCHEMA-MINER pro يقوم بأتمتة أجزاء كبيرة من عملية مواءمة الأنطولوجيا مع الاحتفاظ بفرص للإشراف من قبل الخبراء، مما يحسن من الكفاءة وقابلية التوسع في مهام استخراج المعرفة عبر مجالات علمية متنوعة. بشكل عام، يمثل SCHEMA-MINER pro خطوة كبيرة إلى الأمام في إنشاء إطار عمل شامل من البداية إلى النهاية لهندسة المخططات الدلالية.

Journal: Semantic Web, Volume: 17, Issue: 3
DOI: https://doi.org/10.1177/22104968261431521
Publication Date: 2026-04-25
Author(s): Sameer Sadruddin et al.
Primary Topic: Scientific Computing and Data Management

Overview

The section presents SCHEMA-MINER pro, a human-in-the-loop framework designed to extract and ground structured schemas from scientific literature, addressing the challenge of representing scientific processes in a computationally manageable format. By integrating large language models for schema extraction with an agent-based system, the framework aligns extracted elements to external ontologies through interpretable, multi-step reasoning. The agent employs lexical heuristics, semantic similarity, and expert feedback to ensure the accuracy of the grounding process. The framework was demonstrated on semiconductor manufacturing workflows, specifically atomic layer deposition (ALD) and atomic layer etching (ALE), successfully mapping process parameters and outputs to the QUDT (Quantities, Units, Dimensions, and Types) ontology.

In conclusion, SCHEMA-MINER pro represents a significant advancement in scientific schema discovery and ontology grounding, with implications extending beyond semiconductor manufacturing to the broader Semantic Web. The ontology-grounded schemas produced are machine-actionable and reusable across various scientific domains, facilitating reproducible workflows, automated reasoning, and semantic interoperability. For instance, an ontology-grounded ALD schema can enhance the annotation of experimental datasets and support unit-consistent comparisons. Future directions include aligning schemas with Wikidata to connect process-level knowledge to broader scientific entities and expanding the framework’s application to biology and environmental science, addressing similar issues of unstructured reporting and semantic fragmentation. The integration of SCHEMA-MINER pro with active Linked Open Data (LOD) initiatives and FAIR digital objects aims to contribute to a structured and queryable infrastructure for open science.

Introduction

The introduction of the research paper discusses the critical need for extracting structured information from unstructured scientific literature to enhance the modeling of complex real-world processes. Despite advancements in information extraction (IE) methods, traditional approaches remain labor-intensive and often require extensive labeled datasets, particularly in specialized fields such as materials science and life sciences. Recent developments in large language models (LLMs) have shown promise in automating the extraction of structured data, yet challenges persist due to the specialized language and varying terminologies across publications. To address these issues, the paper introduces SCHEMA-MINER pro, an enhanced framework that integrates agent-based workflows with LLMs for ontology grounding, thereby improving the precision and interpretability of extracted schemas.

SCHEMA-MINER pro employs a systematic, iterative approach that combines heuristic string matching and semantic vector search to align schema elements with ontology concepts, incorporating expert validation throughout the process. This contrasts with traditional monolithic prompting methods, which often lack transparency and require extensive human validation. The framework is evaluated across two semiconductor processes, atomic layer deposition (ALD) and atomic layer etching (ALE), demonstrating its effectiveness in handling complex scientific terminologies. Additionally, a user-friendly web-based interface is introduced to facilitate expert engagement and enhance the accessibility of the schema discovery process. The findings underscore the importance of domain expert involvement in refining schemas, highlighting that a hybrid feedback approach significantly improves the quality and coherence of the extracted information.

Methods

The research employed a tool called SCHEMA-MINER, developed in Python using the LangChain framework, to facilitate schema extraction from various large language models (LLMs), including GPT-4o, GPT-4turbo, and LLaMA 3.1 (8B). The experiments were conducted on a machine with a 16-core CPU and 32 GB of RAM, utilizing cloud services for LLM inferences. The study evaluated schema discovery through four experimental configurations involving different types of expert feedback—descriptive text and direct schema edits—across 21 experiments. Notably, the configuration that combined both feedback types in every iteration yielded the most accurate results for both atomic layer deposition (ALD) and atomic layer etching (ALE) processes.

The significance of the Quantities, Units, Dimensions, and Types (QUDT) ontology in materials science was highlighted, emphasizing its role in standardizing machine-readable representations of physical quantities. This standardization is crucial for ensuring consistency across diverse data sources, enabling automated unit conversion, and enhancing the reproducibility and interoperability of experimental data. The findings indicated that GPT-4turbo outperformed LLaMA 3.1 (8B) in terms of structural coherence, as evidenced by ROUGE-L scores, while maintaining strong semantic alignment across models, as shown by high BERTScores. Overall, GPT-4o and GPT-4turbo demonstrated superior performance in capturing the nuances of ALD and ALE experimental schemas, underscoring their robustness in processing complex scientific data.

Results

In this section, the performance of the SCHEMA-MINER pro framework is evaluated through both quantitative and qualitative analyses, focusing on ALD and ALE processes. The quantitative evaluation assesses the variance and structural differences in schemas generated by three large language models (LLMs): GPT-4o, GPT-4-turbo, and LLaMA 3.1 (8B) across three stages of schema refinement. The evaluation employs three metrics—ROUGE-L, BLEU Score, and BERTScore—to measure schema similarity, with results indicating varying degrees of alignment and evolution among the models. All generated schemas are publicly accessible in the repository.

The qualitative evaluation emphasizes the stability of LLM outputs, which is crucial for maintaining schema coherence during refinement. Domain expert feedback revealed that both GPT-4o and LLaMA 3.1 (8B) exhibited high stability in the ALD process, while GPT-4-turbo showed instability by introducing unvalidated properties. In the ALE process, GPT-4o and GPT-4-turbo maintained relative stability, whereas LLaMA 3.1 (8B) produced several irrelevant properties, particularly in the final stage, negatively impacting the semantic accuracy of the schemas. These findings underscore the importance of model selection in achieving reliable schema extraction.

Discussion

The discussion section of the research paper elaborates on the advancements and contributions of the SCHEMA-MINER pro system, which enhances schema induction and ontology grounding in scientific literature. Building on previous methodologies, SCHEMA-MINER pro integrates large language models (LLMs) with human-in-the-loop workflows to address the complexities of scientific texts, which often contain dense terminology and intricate processes. The paper reviews prior work in schema discovery, highlighting the evolution from rule-based techniques to neural approaches that leverage generative models for schema induction. Notably, it emphasizes the limitations of existing systems, which typically focus on general domains and lack the necessary adaptability for scientific contexts.

The SCHEMA-MINER pro system distinguishes itself by offering a modular architecture that supports agentic AI workflows, allowing for dynamic management of the schema induction process. This includes the incorporation of expert feedback throughout the workflow, enhancing the precision and reusability of the schemas produced. The paper also discusses the integration of ontology grounding, where schema components are aligned with established ontologies to ensure interoperability and facilitate knowledge integration. The agent-based approach employed in SCHEMA-MINER pro automates significant portions of the ontology alignment process while retaining opportunities for expert oversight, thereby improving efficiency and scalability in knowledge extraction tasks across diverse scientific domains. Overall, SCHEMA-MINER pro represents a significant step forward in creating a comprehensive, end-to-end framework for semantic schema engineering.

شارك: