DAPFAM: مجموعة بيانات على مستوى العائلة مدركة للمجال لتقييم استرجاع براءات الاختراع عبر المجالات
DAPFAM: A Domain-Aware Family-level Dataset to benchmark cross domain patent retrieval

شارك:
المجلة: Array، المجلد: 29
DOI: https://doi.org/10.1016/j.array.2026.100720
تاريخ النشر: 2026-02-18
المؤلف: Iliass Ayaou وآخرون
الموضوع الرئيسي: الملكية الفكرية وبراءات الاختراع

نظرة عامة

تقدم البحث DAPFAM، وهو معيار جديد على مستوى العائلة مصمم لتعزيز استرجاع الفن السابق لبراءات الاختراع من خلال معالجة التحديات عبر المجالات بشكل صريح من خلال تقسيمات محددة داخل المجال وخارج المجال بناءً على مخطط تداخل IPC3 جديد. تتكون مجموعة البيانات من 1,247 عائلة استعلام و45,336 عائلة مستهدفة، مما يسهل تقييمًا شاملاً لأنظمة الاسترجاع عبر تكوينات مختلفة، بما في ذلك الطرق المعجمية (BM25) والكثيفة (المحول)، واسترجاع الوثائق والمقاطع، وتقنيات الدمج الهجينة مثل دمج الترتيب المتبادل (RRF). تشير النتائج إلى وجود فجوة كبيرة بين المجالات، حيث كانت أداء خارج المجال أقل بحوالي خمس مرات من داخل المجال عبر جميع التكوينات المختبرة. من الجدير بالذكر أن استرجاع المقاطع يتفوق باستمرار على استرجاع الوثائق، وعلى الرغم من أن الطرق الكثيفة تظهر بعض التحسينات مقارنة بـ BM25، إلا أنها لا تسد الفجوة بين المجالات بشكل كافٍ.

في الختام، يعد DAPFAM بمثابة منصة اختبار منهجية وقابلة للتكرار لتقييم طرق استرجاع معلومات براءات الاختراع عبر المجالات التكنولوجية، مما يكشف عن رؤى حاسمة حول استراتيجيات الاسترجاع وفعاليتها. يبرز الدراسة أن تمثيلات واستراتيجيات الاسترجاع المناسبة ضرورية للأداء القوي، خاصة في مواجهة الانحراف المعجمي والتباينات في أنماط صياغة براءات الاختراع عبر المجالات. بينما تحتوي مجموعة البيانات على معلومات غنية بالبيانات الوصفية وتدعم تطبيقات بحث براءات الاختراع المختلفة، إلا أنها محدودة ببراءات الاختراع باللغة الإنجليزية من Lens.org، مما قد يقدم تحيزات. يعد تقسيم المجالات القائم على IPC3 واحدًا من العديد من الأساليب الممكنة، وقد تحد من قيود التقييم على سلوكيات النماذج رؤى أوسع. بشكل عام، يوفر DAPFAM موردًا قيمًا لتعزيز البحث في استرجاع براءات الاختراع والحقول ذات الصلة.

مقدمة

تتناول مقدمة ورقة البحث الدور الحاسم لاسترجاع الفن السابق لبراءات الاختراع في تعزيز الابتكار واستراتيجية الملكية الفكرية والتقدم التكنولوجي. تبرز التحديات التي تطرحها الزيادة في حجم الابتكارات وطبيعتها متعددة التخصصات، مما يعقد اكتشاف الفن السابق الفعال، خاصة عندما تمتد براءات الاختراع ذات الصلة عبر مجالات تكنولوجية متعددة. لا تلبي المعايير الحالية لاسترجاع براءات الاختراع الاحتياجات في معالجة السيناريوهات عبر المجالات، حيث تنتمي البراءات إلى رموز تصنيف براءات الاختراع الدولية (IPC) أو تصنيف براءات الاختراع التعاوني (CPC) المختلفة، مما يؤدي إلى فجوات كبيرة في فعالية الاسترجاع.

تؤكد الورقة على قيود مجموعات بيانات استرجاع براءات الاختراع الحالية، التي غالبًا ما تركز على ولايات قضائية واحدة وتفتقر إلى تقسيم منهجي للمجالات لتقييم عبر المجالات. تعيق هذه القصور المقارنات الدقيقة لطرق الاسترجاع تحت ظروف مجالات مختلفة. يقترح المؤلفون نهجًا جديدًا، DAPFAM، الذي يقسم المجالات بشكل منهجي بناءً على تداخل تصنيف IPC3، مما يسمح بتقييم أوضح لتحديات الاسترجاع عبر المجالات. يقدمون أدلة تجريبية تظهر أن حوالي 26% من الوثائق ذات الصلة لا تشارك رموز IPC مع استعلاماتهم، مما يبرز أهمية معالجة مشكلات الاسترجاع خارج المجال. تشير النتائج إلى أن طرق الاسترجاع التقليدية، وخاصة النماذج الكثيفة، تعاني من تدهور كبير في الأداء في السياقات عبر المجالات، مما يبرز الحاجة إلى تقنيات تكيف المجال المتخصصة لتعزيز فعالية الاسترجاع عبر المناظر التكنولوجية المتنوعة.

طرق

في هذا القسم، يناقش المؤلفون تطبيق الطرق العصبية وتمثيلات النصوص الدلالية في استرجاع معلومات براءات الاختراع (IR)، مع تسليط الضوء على الانتقال من الطرق التقليدية مثل كيس الكلمات وn-gram إلى التضمينات العصبية المتقدمة. ساعدت تقنيات مثل Word2Vec وDoc2Vec في تطوير تضمينات خاصة بالبراءات التي تسد الفجوة المعجمية من خلال توليد تمثيلات متجهة تركز على المجال. أظهر استخدام النماذج المعتمدة على المحولات، وخاصة BERT ونسخها المعدلة حسب المجال، وعدًا في تعزيز الفهم الدلالي السياقي داخل مجموعات براءات الاختراع، مما أدى إلى تحسين المطابقة على مستوى العبارات والتعامل بشكل أفضل مع هياكل المطالبات الهرمية.

يستكشف المؤلفون أيضًا تقنيات الهجين، وخاصة دمج الترتيب المتبادل (RRF)، الذي يعزز فعالية الاسترجاع عبر أنظمة تشغيل مختلفة، بما في ذلك الدمج القابل للمقاطع والوثائق فقط. يوضحون تنفيذ استراتيجيات استرجاع على مستوى المقاطع، مثل maxP وavgP وsumP، لتجميع الأدلة على مستوى الوثيقة. تشير نتائجهم إلى أن معلمة مثلى أعلى $K=60$ تحسن تشتت الترتيب في الواجهات الخلفية على مستوى الوثيقة، بينما يقدم الدمج الهجين للوثائق فقط استراتيجية فعالة من حيث الموارد توازن بين الفعالية والتكلفة الحسابية، مما يجعلها مناسبة للأنظمة الإنتاجية ذات قيود الميزانية.

نتائج

يستعرض قسم النتائج نتائج تجريبية واسعة مستمدة من 249 تكوينًا فريدًا، مع التركيز على عوامل مختلفة تؤثر على فعالية الاسترجاع. يتم هيكلة التحليل حول ستة مجالات رئيسية: (1) مقارنات الأداء عبر الواجهات الخلفية الفردية، (2) فعالية تمثيلات الاستعلام المختلفة، (3) تحسين أطوال المقاطع، (4) تقييم استراتيجيات التجميع، (5) أداء طرق الدمج الهجينة، و(6) التوازن بين الكفاءة والفعالية. جميع المقاييس المبلغ عنها، بما في ذلك المكسب التراكمي المخفض المنظم عند الترتيب 100 (NDCG@100) والاسترجاع عند الترتيب 100، يتم حسابها كمتوسط عبر الاستعلامات ضمن مجموعات التقييم المحددة، مما يوفر نظرة شاملة على النتائج التجريبية.

نقاش

في هذا القسم، يقدم المؤلفون DAPFAM، وهو معيار جديد لاسترجاع براءات الاختراع على مستوى العائلة مصمم لتعزيز تقييم طرق الاسترجاع عبر المجالات. يعالج DAPFAM الفجوات الموجودة في مجموعات بيانات استرجاع براءات الاختراع من خلال توفير تقسيمات تقييم صريحة داخل المجال وخارج المجال بناءً على رموز تصنيف براءات الاختراع الدولية (IPC). يسمح هذا التمييز المنهجي بتقييم مباشر لتحديات الاسترجاع عبر المجالات، وهو أمر حاسم لتحسين أنظمة البحث عن براءات الاختراع. تجمع مجموعة البيانات براءات الاختراع على مستوى العائلة، مما يوازن الاستعلامات عبر مجالات تكنولوجية مختلفة، وبالتالي تسهيل البحث القابل للتكرار في استراتيجيات الاسترجاع المدركة للمجال وتقنيات الدمج الهجينة.

يبرز المؤلفون ثلاث مساهمات كبيرة لعملهم: تقديم أول معيار لاسترجاع براءات الاختراع على مستوى العائلة مع تقييم صريح خارج المجال، وتحليل تجريبي شامل للاختيارات التصميمية التي تؤثر على أداء الاسترجاع، وتحديد تحدٍ أساسي في استرجاع خارج المجال حيث تعاني الطرق الكثيفة من أداء أقل مقارنة بالطرق المعجمية. تشير النتائج إلى أن الأداء يمكن أن ينخفض بحوالي خمس مرات في السيناريوهات خارج المجال، مما يبرز الحاجة إلى استراتيجيات فعالة في اكتشاف الفن السابق عبر المجالات. تتوفر مجموعة بيانات DAPFAM للبحث، كاملة مع بيانات وصفية وأحكام ملاءمة، تدعم تكوينات تجريبية متنوعة وتضمن القابلية للتكرار في الدراسات المستقبلية.

Journal: Array, Volume: 29
DOI: https://doi.org/10.1016/j.array.2026.100720
Publication Date: 2026-02-18
Author(s): Iliass Ayaou et al.
Primary Topic: Intellectual Property and Patents

Overview

The research introduces DAPFAM, a novel family-level benchmark designed to enhance patent prior-art retrieval by explicitly addressing cross-domain challenges through defined IN-domain and OUT-domain partitions based on a new IPC3 overlap scheme. The dataset comprises 1,247 query families and 45,336 target families, facilitating a comprehensive evaluation of retrieval systems across various configurations, including lexical (BM25) and dense (transformer) methods, document and passage-level retrieval, and hybrid fusion techniques like Reciprocal Rank Fusion (RRF). The findings indicate a significant domain gap, with OUT-domain performance being approximately five times lower than IN-domain across all tested configurations. Notably, passage-level retrieval consistently outperforms document-level retrieval, and while dense methods show some improvement over BM25, they do not sufficiently bridge the OUT-domain gap.

In conclusion, DAPFAM serves as a systematic and reproducible testbed for evaluating patent information retrieval methods across technological domains, revealing critical insights into retrieval strategies and their effectiveness. The study highlights that appropriate retrieval representations and strategies are crucial for robust performance, particularly in the face of lexical drift and variations in patent drafting styles across domains. While the dataset is rich in metadata and supports various patent research applications, it is limited to English-language patents from Lens.org, which may introduce biases. The IPC3-based domain partitioning is one of many possible approaches, and the evaluation’s constraints on model-specific behaviors may limit broader insights. Overall, DAPFAM provides a valuable resource for advancing research in patent retrieval and related fields.

Introduction

The introduction of the research paper addresses the critical role of patent prior art retrieval in fostering innovation, intellectual property strategy, and technological advancement. It highlights the challenges posed by the increasing volume and interdisciplinary nature of modern innovations, which complicate effective prior art discovery, especially when relevant patents span multiple technological domains. Current patent retrieval benchmarks fall short in addressing cross-domain scenarios, where patents belong to different International Patent Classification (IPC) or Cooperative Patent Classification (CPC) codes, leading to significant gaps in retrieval effectiveness.

The paper emphasizes the limitations of existing patent retrieval datasets, which often focus on single jurisdictions and lack systematic domain partitioning for cross-domain evaluation. This inadequacy hampers rigorous comparisons of retrieval methods under varying domain conditions. The authors propose a novel approach, DAPFAM, which systematically partitions domains based on IPC3 classification overlap, allowing for a clearer evaluation of cross-domain retrieval challenges. They present empirical evidence showing that approximately 26% of relevant documents do not share IPC codes with their queries, underscoring the importance of addressing out-of-domain retrieval issues. The findings indicate that traditional retrieval methods, particularly dense models, experience significant performance degradation in cross-domain contexts, highlighting the need for specialized domain adaptation techniques to enhance retrieval effectiveness across diverse technological landscapes.

Methods

In this section, the authors discuss the application of neural methods and semantic text representations in patent information retrieval (IR), highlighting the transition from traditional bag-of-words and n-gram approaches to advanced neural embeddings. Techniques such as Word2Vec and Doc2Vec have facilitated the development of patent-specific embeddings that effectively bridge the lexical gap by generating domain-centric vector representations. The use of transformer-based models, particularly BERT and its domain-adapted variants, has shown promise in enhancing contextual semantic understanding within patent corpora, leading to improved phrase-level matching and better handling of hierarchical claim structures.

The authors also explore hybridization techniques, specifically reciprocal rank fusion (RRF), which enhances retrieval effectiveness across different operational regimes, including passage-capable and document-only fusion. They detail the implementation of passage-level retrieval strategies, such as maxP, avgP, and sumP, to consolidate evidence at the document level. Their findings indicate that a higher optimal parameter $K=60$ improves rank dispersion in document-level backends, while document-only hybrid fusion offers a resource-efficient strategy that balances effectiveness and computational cost, making it suitable for production systems with budget constraints.

Results

The results section details extensive experimental findings derived from 249 unique configurations, focusing on various factors influencing retrieval effectiveness. The analysis is structured around six key areas: (1) comparisons of performance across single backends, (2) effectiveness of different query representations, (3) optimization of passage lengths, (4) evaluation of aggregation strategies, (5) performance of hybrid fusion methods, and (6) trade-offs between efficiency and effectiveness. All reported metrics, including Normalized Discounted Cumulative Gain at rank 100 (NDCG@100) and Recall at rank 100, are averaged across queries within specified evaluation subsets, providing a comprehensive overview of the experimental outcomes.

Discussion

In this section, the authors present DAPFAM, a novel family-level patent retrieval benchmark designed to enhance the evaluation of cross-domain retrieval methods. DAPFAM addresses existing gaps in patent retrieval datasets by providing explicit IN-domain and OUT-domain evaluation partitions based on International Patent Classification (IPC) codes. This systematic distinction allows for a direct assessment of cross-domain retrieval challenges, which is critical for improving patent search systems. The dataset aggregates patents at the family level, balancing queries across various technological domains, thus facilitating reproducible research into domain-aware retrieval strategies and hybrid fusion techniques.

The authors highlight three significant contributions of their work: the introduction of the first family-level patent benchmark with explicit out-of-domain evaluation, a comprehensive empirical analysis of design choices affecting retrieval performance, and the identification of a fundamental challenge in OUT-domain retrieval where dense methods underperform compared to lexical approaches. The findings indicate that performance can drop by approximately five times in OUT-domain scenarios, underscoring the need for effective strategies in cross-domain prior art discovery. The DAPFAM dataset is made available for research, complete with metadata and relevance judgments, supporting various experimental configurations and ensuring reproducibility in future studies.

شارك: