DOI: https://doi.org/10.14778/3797919.3797940
تاريخ النشر: 2026-02-01
المؤلف: Xinyi Zhang وآخرون
الموضوع الرئيسي: أنظمة قواعد البيانات المتقدمة والاستعلامات
نظرة عامة
يقدم القسم SysInsight، وهو نظام جديد لضبط قواعد البيانات مدفوع بالشيفرة مصمم لتحسين أداء أنظمة إدارة قواعد البيانات الحديثة (DBMSs) من خلال استخراج معرفة ضبط دقيقة مباشرة من شيفرة المصدر الخاصة بـ DBMS. تعتمد طرق الضبط الآلي التقليدية إما على أساليب مدفوعة بالبيانات، والتي يمكن أن تكون مكلفة من حيث الموارد، أو على heuristics يدوية تفتقر إلى التحديد. يتناول SysInsight هذه القيود من خلال استخدام تحليل الشيفرة الثابت واستدلالات قائمة على نماذج اللغة الكبيرة (LLM) لتحديد مسارات التنفيذ المتأثرة بعناصر التكوين واستنباط رؤى ضبط دلالية. يتم تحويل هذه الرؤى بعد ذلك إلى قواعد ضبط كمية من خلال استخراج قواعد الارتباط، مما يسهل عملية ضبط أكثر كفاءة.
تشير النتائج إلى أن SysInsight يعزز بشكل كبير من كفاءة الضبط، حيث يحقق التقارب إلى التكوينات المثلى بمعدل أسرع بمعدل 7.11 مرة من الطرق الحديثة (SOTA)، بينما يقدم أيضًا تحسينًا بنسبة 19.9% في الأداء. يؤكد المؤلفون أن SysInsight يمثل خطوة نحو الضبط الشفاف من خلال الاستفادة من المعرفة الداخلية للنظام، ويعترفون بإمكانية دمج مصادر بيانات إضافية، مثل الوثائق وبيانات الأداء التاريخية، لتحسين استراتيجيات الضبط بشكل أكبر. ستركز الأعمال المستقبلية على تطوير أساليب هجينة تجمع بين هذه المصادر المتنوعة لتعزيز قابلية التكيف وفعالية وكلاء الضبط.
مقدمة
تناقش مقدمة الورقة التحديات المرتبطة بتكوين أنظمة إدارة قواعد البيانات الحديثة (DBMSs)، والتي تقدم العديد من عناصر التكوين الأساسية لتحسين الأداء. تُصنف مهمة تحديد الإعدادات المثلى على أنها NP-hard، مما يدفع مديري قواعد البيانات (DBAs) إلى استثمار وقت كبير في الضبط اليدوي. لمعالجة هذه المشكلة، استكشفت الأبحاث الحديثة أتمتة عملية الضبط من خلال تقنيات التعلم الآلي (ML)، والتي تُصنف إلى نموذجين رئيسيين: طرق مدفوعة بالبيانات وطرق مدفوعة يدوياً. تعالج الأساليب المدفوعة بالبيانات الضبط كمشكلة تحسين مغلقة، معتمدة على جمع بيانات الأداء بشكل تكراري، بينما تستخدم الأساليب المدفوعة يدوياً نماذج لغوية مسبقة التدريب لاستخراج heuristics الضبط من الوثائق. ومع ذلك، فإن فعالية الأساليب المدفوعة يدوياً محدودة بتوافر التوصيات الخبيرة الموثقة، حيث تشير الدراسات إلى أن أقل من 50% من العناصر لديها نطاقات ضبط ذات مغزى.
يقدم القسم أيضًا SysInsight، وهو نظام مصمم لتعزيز عملية الضبط من خلال استرجاع القواعد ذات الصلة بناءً على سياق التشغيل الحالي وتحديث إحصائياتها باستمرار للحفاظ على الموثوقية. يقوم SysInsight بتحديد القواعد المرشحة من خلال مطابقة الشروط السابقة ويصنفها وفقًا لتحسينها المتوقع، والذي يُعرف بأنه ناتج متوسط زيادة الأداء واحتمالية تحقيق تلك الزيادة. بعد تطبيق التكوينات المقترحة، يقوم SysInsight بتحديث إحصائيات القواعد ذات الصلة، مع التركيز بشكل خاص على تلك التي تتماشى تعديلاتُها مع التغييرات الفعلية المطبقة. يسمح هذا التحديث الديناميكي بإعادة معايرة التحسين المتوقع للقواعد بناءً على ملاحظات الأداء الجديدة، مما يُحسن التوجيه الضبطي المقدم إلى DBAs.
طرق
في قسم الطرق، يقارن البحث بين أساليب مختلفة لضبط نظام إدارة قواعد البيانات (DBMS)، مع التركيز على أداء GPTuner وSysInsight وطرق مدفوعة بالبيانات الأخرى. يستخدم GPTuner نماذج لغوية كبيرة (LLMs) لاستخراج المعرفة الهيكلية من الوثائق، بما في ذلك النطاقات والقيم المقترحة لعناصر الضبط. ومع ذلك، فإنه يواجه قيودًا بسبب الطبيعة الخشنة للمعلومات المتاحة، حيث تقدم 20 من أصل 44 عنصرًا قيمًا مميزة للحد الأدنى أو الأقصى، ويفتقر العديد منها إلى توجيهات محددة للسياق. في المقابل، يظهر SysInsight، الذي يستخرج قواعد الضبط مباشرة من الشيفرة المصدرية، أداءً متفوقًا من خلال توفير معلومات تحكم دقيقة وتوصيات مدركة للسياق، حيث ينجح في تحديد قواعد الضبط لـ 33 عنصرًا.
يمتد المقارنة إلى الأساليب المدفوعة بالبيانات، حيث يُلاحظ أن SMAC وDDPG++ غير قادرين على العثور على التكوينات المثلى خلال 20 تكرارًا بسبب اعتمادهما على البدء من الصفر. بينما تستفيد OtterTune وResTune من البيانات التاريخية، فإن فعاليتها تتناقص عندما تختلف أحمال العمل المستهدفة بشكل كبير عن أحمال العمل المصدر. ومع ذلك، يلتقط SysInsight العلاقات القابلة للتعميم بين سلوكيات النظام وتعديلات العناصر، مما يمكّن من تحسين قوي عبر سياقات متنوعة. بالإضافة إلى ذلك، يقيم البحث طرق اختيار العناصر المختلفة، كاشفًا أن النهج القائم على الشيفرة يتفوق على الأساليب القائمة على LLM والأساليب القائمة على ML من خلال تحديد دقيق لوظائف الاختناق والعناصر المرتبطة بها، بينما تكافح الطريقة القائمة على ML بسبب عدم كفاية الملاحظات لتصنيفات الأهمية الموثوقة.
نقاش
في هذا القسم، يناقش المؤلفون التحديات والابتكارات في ضبط تكوين قواعد البيانات، وخاصة من خلال نهج جديد يُسمى SysInsight. يبرزون قيود الأساليب اليدوية والمدفوعة بالبيانات الحالية، التي غالبًا ما تعتمد على إرشادات مبسطة للغاية وتفشل في تقديم توصيات مخصصة لأحمال العمل المحددة. يجادل المؤلفون بأن نموذج الضبط المدفوع بالشيفرة، الذي يستخرج الرؤى مباشرة من شيفرة المصدر لنظام قاعدة البيانات، يمكن أن ينتج استراتيجيات ضبط أكثر دقة وقابلية للتنفيذ. يقوم SysInsight بتحديد العلاقات السببية بين عناصر التكوين وسلوك النظام، مما يمكّنه من صياغة فرضيات وتحسينها إلى قواعد ضبط موثوقة من خلال التحقق التجريبي.
يُفصل المؤلفون التعقيدات المرتبطة بالضبط الآلي المدفوع بالشيفرة، بما في ذلك الحاجة إلى تحليل فعال لقاعدة الشيفرة الكبيرة والتحدي المتمثل في تحويل الرؤى إلى توجيهات موثوقة. يقترحون إطار عمل يجمع بين التحليل الثابت مع قدرات نماذج اللغة الكبيرة (LLM) لاستخراج معرفة الضبط، وتحديد اختناقات الأداء، وتوليد قواعد ضبط مدركة للسياق. لا يعزز هذا الإطار كفاءة عملية الضبط فحسب، بل يضمن أيضًا أن تكون التوصيات مستندة إلى سلوك النظام الفعلي، مما يعالج الحاجة الملحة للموثوقية في بيئات الضبط ذات المخاطر العالية. يتم إثبات فعالية SysInsight من خلال تجارب واسعة، تُظهر تحسينًا كبيرًا في سرعة الضبط والأداء مقارنة بالطرق الحالية.
DOI: https://doi.org/10.14778/3797919.3797940
Publication Date: 2026-02-01
Author(s): Xinyi Zhang et al.
Primary Topic: Advanced Database Systems and Queries
Overview
The section presents SysInsight, a novel code-driven database tuning system designed to optimize the performance of modern database management systems (DBMSs) by extracting fine-grained tuning knowledge directly from the DBMS source code. Traditional automated tuning methods either rely on data-driven approaches, which can be resource-intensive, or on manual heuristics that lack specificity. SysInsight addresses these limitations by employing static code analysis and large language model (LLM)-based reasoning to identify execution paths influenced by configuration knobs and to derive semantic tuning insights. These insights are subsequently transformed into quantitative tuning rules through association rule mining, facilitating a more efficient tuning process.
The results indicate that SysInsight significantly enhances tuning efficiency, achieving convergence to optimal configurations an average of 7.11 times faster than state-of-the-art (SOTA) methods, while also delivering a 19.9% improvement in performance. The authors emphasize that SysInsight represents a step towards white-box tuning by leveraging internal system knowledge, and they acknowledge the potential for integrating additional data sources, such as documentation and historical performance data, to further improve tuning strategies. Future work will focus on developing hybrid approaches that combine these diverse sources to enhance the adaptability and effectiveness of tuning agents.
Introduction
The introduction of the paper discusses the challenges associated with configuring modern database management systems (DBMSs), which present numerous configuration knobs essential for optimizing performance. The task of identifying optimal settings is classified as NP-hard, leading database administrators (DBAs) to invest considerable time in manual tuning. To address this issue, recent research has explored automating the tuning process through machine learning (ML) techniques, which are categorized into two main paradigms: data-driven methods and manual-driven methods. Data-driven approaches treat tuning as a black-box optimization problem, relying on iterative performance data collection, while manual-driven methods utilize pretrained language models to extract tuning heuristics from documentation. However, the effectiveness of manual-driven methods is limited by the availability of documented expert recommendations, with studies indicating that less than 50% of knobs have meaningful tuning ranges.
The section also introduces SysInsight, a system designed to enhance the tuning process by retrieving relevant rules based on the current runtime context and continuously updating their statistics to maintain reliability. SysInsight identifies candidate rules by matching antecedent predicates and ranks them according to their expected improvement, defined as the product of average performance gain and the probability of achieving that gain. Following the application of suggested configurations, SysInsight updates the statistics of relevant rules, particularly focusing on those whose adjustments align with actual applied changes. This dynamic updating process allows for recalibrating the expected improvement of rules based on new performance observations, thereby refining the tuning guidance provided to DBAs.
Methods
In the methods section, the research compares various approaches for database management system (DBMS) tuning, focusing on the performance of GPTuner, SysInsight, and other data-driven methods. GPTuner utilizes large language models (LLMs) to extract structured knowledge from documentation, including suggested ranges and values for tuning knobs. However, it faces limitations due to the coarse-grained nature of the information available, with only 20 out of 44 knobs offering distinct min or max values, and many lacking context-specific guidance. In contrast, SysInsight, which derives tuning rules directly from source code, demonstrates superior performance by providing fine-grained control information and context-aware recommendations, successfully identifying tuning rules for 33 knobs.
The comparison extends to data-driven methods, where SMAC and DDPG++ are noted for their inability to find optimal configurations within 20 iterations due to their reliance on starting from scratch. While OtterTune and ResTune leverage historical data, their effectiveness diminishes when target workloads differ significantly from source workloads. SysInsight, however, captures generalizable relationships between system behaviors and knob adjustments, enabling robust optimization across diverse contexts. Additionally, the study evaluates different knob selection methods, revealing that the code-based approach outperforms LLM-based and ML-based methods by accurately identifying bottleneck functions and associated knobs, while the ML-based method struggles due to insufficient observations for reliable importance rankings.
Discussion
In this section, the authors discuss the challenges and innovations in database configuration tuning, particularly through a novel approach termed SysInsight. They highlight the limitations of existing manual and data-driven tuning methods, which often rely on oversimplified guidelines and fail to provide tailored recommendations for specific workloads. The authors argue that a code-driven tuning paradigm, which extracts insights directly from the database system’s source code, can yield more precise and actionable tuning strategies. SysInsight identifies causal relationships between configuration knobs and system behavior, enabling it to formulate hypotheses and refine them into reliable tuning rules through empirical validation.
The authors detail the complexities involved in automatic code-driven tuning, including the need for effective analysis of large codebases and the challenge of translating insights into reliable guidance. They propose a framework that combines static analysis with large language model (LLM) capabilities to extract tuning knowledge, identify performance bottlenecks, and generate context-aware tuning rules. This framework not only enhances the efficiency of the tuning process but also ensures that recommendations are grounded in actual system behavior, thereby addressing the critical need for robustness in high-stakes tuning environments. The effectiveness of SysInsight is demonstrated through extensive experiments, showing a significant improvement in tuning speed and performance compared to existing methods.
