كشف الشذوذ لمراقبة جودة البيانات الآلية في كاشف CMS
Anomaly Detection for Automated Data Quality Monitoring in the CMS Detector

شارك:
المجلة: EPJ Research Infrastructures، المجلد: 10، العدد: 1
DOI: https://doi.org/10.1007/s41781-025-00147-2
PMID: https://pubmed.ncbi.nlm.nih.gov/41726844
تاريخ النشر: 2026-02-09
المؤلف: A. Brinkerhoff وآخرون
الموضوع الرئيسي: دراسات فيزياء الجسيمات النظرية والتجريبية

نظرة عامة

يقدم هذا القسم نظام “AutoDQM”، المصمم لمراقبة جودة البيانات تلقائيًا في كاشفات الجسيمات الكبيرة مثل الكاشف المدمج للميونات (CMS) في مصادم الهادرونات الكبير (CERN). يستخدم هذا النظام طرقًا إحصائية متقدمة وتقنيات تعلم آلي غير خاضعة للإشراف لتسهيل تقييمات جودة البيانات بسرعة وشمولية. بشكل محدد، يستخدم خوارزميات كشف الشذوذ التي تتضمن دالة الاحتمال بيتا-ثنائية الحد، وتحليل المكونات الرئيسية، وتقييم صور مشفرات الشبكة العصبية.

تظهر فعالية AutoDQM من خلال قدرته على تحديد البيانات الشاذة “السيئة” – البيانات التي تأثرت بشكل كبير بأعطال الكاشف – بمعدل أعلى من 4 إلى 6 مرات مقارنة بالبيانات “الجيدة”. تؤكد هذه النتيجة على إمكانية استخدام AutoDQM كأداة قوية لمراقبة جودة البيانات بشكل عام في تجارب الفيزياء عالية الطاقة، مما يضمن سلامة البيانات المجمعة خلال تصادمات البروتون-بروتون في عام 2022.

مقدمة

تجربة الكاشف المدمج للميونات (CMS) هي كاشف جسيمات متطور في مصادم الهادرونات الكبير (LHC) في CERN، يهدف بشكل أساسي إلى تحليل تصادمات البروتون-بروتون عالية الطاقة. لعبت دورًا حاسمًا في الاكتشاف المشترك لهيغز بوسون جنبًا إلى جنب مع تجربة ATLAS، وتركز حاليًا على استكشاف ظواهر فيزيائية جديدة، بما في ذلك المادة المظلمة وعدم التماثل بين المادة والمادة المضادة في الكون. يستخدم CMS خوارزمية “تدفق الجسيمات” الشاملة لتحديد وقياس الجسيمات المختلفة، مستفيدًا من جهاز تتبع سيليكون متعدد الطبقات وكالوريمترات متقدمة لالتقاط وتحليل بيانات التصادم. نظرًا لمعدلات التصادم العالية في LHC، فإن مراقبة جودة البيانات (DQM) الفعالة ضرورية لضمان سلامة البيانات المجمعة، مما يتطلب معالجة وتحليل في الوقت الحقيقي لتحديد الشذوذ.

لزيادة كفاءة DQM، يقدم البحث AutoDQM، أداة قائمة على الويب تستفيد من الطرق الإحصائية وتقنيات التعلم الآلي غير الخاضعة للإشراف للمراقبة التلقائية. تستخدم هذه الأداة تقنيات مثل دالة الاحتمال بيتا-ثنائية الحد، وتحليل المكونات الرئيسية (PCA)، ومشفرات الشبكة العصبية لكشف الشذوذ. يتم مناقشة دراسات الأداء لـ AutoDQM، خاصة باستخدام مخططات مراقبة مستوى 1 من مجموعة بيانات 2022، مما يبرز تطبيقها في مراقبة كاشف الميونات. تؤكد النتائج على أهمية تطوير أدوات قوية مثل AutoDQM لتبسيط تحديد المشكلات داخل كاشف CMS المعقد، مما يحسن في النهاية جودة البيانات ويسهل البحث الحساس عن فيزياء جديدة.

النتائج

تشير النتائج إلى أن كل من خوارزميات بيتا-ثنائية الحد وPCA تميز بشكل فعال بين الجولات الجيدة والسيئة في مجموعة البيانات. بشكل محدد، عندما يكون متوسط عدد علامات المدرج البياني منخفضًا (أقل من 3) للجولات الجيدة، تكشف مخططات ROC عالية التردد (HF) أن الجولات السيئة تظهر 3 إلى 4 مرات أكثر من العلامات. تؤكد مخططات ROC للغابة العشوائية (RF) هذه النتيجة، حيث تظهر أنه بينما يتم وضع علامة على أقل من 12% من الجولات الجيدة، تتجاوز 35-50% من الجولات السيئة العتبة. يُعترف بأن AutoDQM، مثل طرق كشف الشذوذ الأخرى، لا يمكنها تحديد جميع الجولات السيئة، حيث قد لا تؤثر بعض المشكلات على مستوى 1 Trigger (L1T) أو قد لا تكون مرئية في المدرجات البيانية DQM على الإنترنت. ومع ذلك، نجح AutoDQM في اكتشاف نصف المشكلات الخطيرة التي تؤثر على جودة بيانات CMS في عام 2022، مع معدل إيجابي زائف منخفض أقل من 12% للجولات الجيدة.

تحسن أداء اختبارات بيتا-ثنائية الحد χ² وZ’ max بشكل كبير مع زيادة عدد الجولات المرجعية، حيث تكون توزيعات شغل الجسيمات المعاد بناؤها حساسة لظروف التداخل المتغيرة. يضمن استخدام عدة جولات مرجعية أن واحدة على الأقل من الجولات ستتطابق على الأرجح مع ظروف التداخل لجولة البيانات التي يتم تقييمها. تأخذ الخوارزميات (AE وPCA) في الاعتبار بشكل طبيعي تباين التداخل، حيث تم تدريبها على مجموعة متنوعة من الجولات. بينما يوجد بعض التباين في الأداء بين الخوارزميات، لا تبرز أي طريقة واحدة كأفضل. يتم تحقيق النتائج المثلى عندما يتم تطبيق جميع اختبارات الجودة الثلاثة في وقت واحد، مما يؤدي إلى مخطط HF ROC الذي يشير إلى 4-6 مرات أكثر من العلامات في الجولات السيئة مقارنة بالجولات الجيدة، مع أكثر من 55% من الجولات السيئة تم وضع علامة عليها مع 3 شذوذات على الأقل، بينما تصل فقط 13% من الجولات الجيدة إلى هذه العتبة. في هذا النهج المشترك، يتم حساب العلامات المميزة، مما يعني أنه إذا تم وضع علامة على المدرج البياني بواسطة اختبارات متعددة، فإن كل حالة تساهم في إجمالي عدد الشذوذات.

المناقشة

تعزز أداة AutoDQM عملية مراقبة جودة البيانات (DQM) في تجربة CMS من خلال أتمتة كشف الشذوذات في المدرجات البيانية الناتجة عن بيانات الكاشف. تعتمد DQM التقليدية على الفحص اليدوي للعديد من المدرجات البيانية، مما يتطلب جهدًا كبيرًا وعرضة للأخطاء. تستخدم AutoDQM اختبارات إحصائية، مثل دالة بيتا-ثنائية الحد، وتقنيات تعلم الآلة، بما في ذلك تحليل المكونات الرئيسية (PCA) ومشفرات الشبكة العصبية (AEs)، لتحديد الانحرافات عن أنماط البيانات المتوقعة. من خلال حساب نسب الاحتمالات وقيم السحب، تبرز AutoDQM بشكل فعال المناطق الشاذة في المدرجات البيانية، مما يسمح للمراقبين بتحديد المشكلات بسرعة التي قد تؤثر على جودة البيانات والتحليلات الفيزيائية اللاحقة.

تم تقييم أداء AutoDQM باستخدام مجموعة بيانات تتكون من 265 جولة جيدة و43 جولة سيئة، حيث نجحت الأداة في تحديد أكثر من 50% من الجولات السيئة بينما وضعت علامة على أقل من 15% من الجولات الجيدة كشاذة. يسمح دمج PCA وAE بكشف شذوذ قوي دون الحاجة إلى بيانات تدريب مصنفة، مما يجعل النظام قابلاً للتكيف مع أنواع مختلفة من الشذوذ. يقلل مقياس كاي-تربيع المعدل المستخدم في AutoDQM من التحيزات المتعلقة بشغل المدرج البياني، مما يضمن أن يبقى كشف الشذوذ فعالًا عبر جولات بأطوال متفاوتة. بشكل عام، تمثل AutoDQM تقدمًا كبيرًا في DQM الآلي، مما يسهل التعرف السريع على مشكلات الكاشف ويعزز موثوقية البيانات المجمعة لأبحاث الفيزياء.

Journal: EPJ Research Infrastructures, Volume: 10, Issue: 1
DOI: https://doi.org/10.1007/s41781-025-00147-2
PMID: https://pubmed.ncbi.nlm.nih.gov/41726844
Publication Date: 2026-02-09
Author(s): A. Brinkerhoff et al.
Primary Topic: Particle physics theoretical and experimental studies

Overview

The section presents the “AutoDQM” system, designed for Automated Data Quality Monitoring in large particle detectors such as the Compact Muon Solenoid (CMS) at the CERN Large Hadron Collider. This system employs advanced statistical methods and unsupervised machine learning techniques to facilitate rapid and thorough data quality assessments. Specifically, it utilizes anomaly detection algorithms that incorporate the beta-binomial probability function, principal component analysis, and neural network autoencoder image evaluation.

The effectiveness of AutoDQM is evidenced by its ability to identify anomalous “bad” data—data significantly impacted by detector malfunctions—at a rate 4 to 6 times higher than that of “good” data. This finding underscores AutoDQM’s potential as a robust tool for general data quality monitoring in high-energy physics experiments, ensuring the integrity of the data collected during proton-proton collisions in 2022.

Introduction

The Compact Muon Solenoid (CMS) experiment is a sophisticated particle detector at the CERN Large Hadron Collider (LHC), primarily aimed at analyzing high-energy proton-proton collisions. It played a crucial role in the joint discovery of the Higgs boson alongside the ATLAS experiment and is currently focused on exploring new physics phenomena, including dark matter and the matter-antimatter asymmetry of the universe. The CMS employs a comprehensive “particle-flow” algorithm to identify and measure various particles, utilizing a multi-layer silicon tracker and advanced calorimeters to capture and analyze collision data. Given the high collision rates at the LHC, effective data quality monitoring (DQM) is essential to ensure the integrity of the collected data, which involves real-time processing and analysis to identify anomalies.

To enhance DQM efficiency, the paper introduces AutoDQM, a web-based tool that leverages statistical methods and unsupervised machine learning for automated monitoring. This tool employs techniques such as the beta-binomial probability function, principal component analysis (PCA), and neural network autoencoders for anomaly detection. Performance studies of AutoDQM, particularly using Level-1 Trigger monitoring plots from the 2022 dataset, are discussed, highlighting its application in muon detector monitoring. The findings underscore the importance of developing robust tools like AutoDQM to streamline the identification of issues within the complex CMS detector, ultimately improving data quality and facilitating sensitive searches for new physics.

Results

The results indicate that both the beta-binomial and PCA algorithms effectively distinguish between good and bad runs in the data set. Specifically, when the mean number of histogram flags is low (less than 3) for good runs, the high-frequency (HF) ROC plots reveal that bad runs exhibit 3 to 4 times more flags. The random forest (RF) ROC plots corroborate this finding, showing that while less than 12% of good runs are flagged, 35-50% of bad runs exceed the threshold. It is acknowledged that AutoDQM, like other anomaly detection methods, cannot identify all bad runs, as some issues may not affect the Level 1 Trigger (L1T) or may not be visible in the online DQM histograms. Nonetheless, AutoDQM successfully detected half of the serious issues impacting CMS data quality in 2022, with a low false positive rate of less than 12% for good runs.

The performance of the beta-binomial χ² and Z’ max tests improves significantly with an increased number of reference runs, as the reconstructed particle occupancy distributions are sensitive to varying pileup conditions. Utilizing multiple reference runs ensures that at least one run will likely match the pileup conditions of the data run being evaluated. The algorithms (AE and PCA) inherently account for pileup variability, having been trained on a diverse set of runs. While there is some variation in performance among the algorithms, no single method stands out as superior. The optimal results are achieved when all three quality tests are applied concurrently, leading to the HF ROC plot indicating 4-6 times more flags in bad runs compared to good runs, with over 55% of bad runs flagged with at least 3 anomalies, while only 13% of good runs reach this threshold. In this combined approach, distinct flags are counted, meaning that if a histogram is flagged by multiple tests, each instance contributes to the total anomaly count.

Discussion

The AutoDQM tool enhances the Data Quality Monitoring (DQM) process in the CMS experiment by automating the detection of anomalies in histograms generated from detector data. Traditional DQM relies on manual inspection of numerous histograms, which is both labor-intensive and prone to error. AutoDQM employs statistical tests, such as the beta-binomial function, and machine learning techniques, including Principal Component Analysis (PCA) and autoencoders (AEs), to identify deviations from expected data patterns. By calculating likelihood ratios and pull values, AutoDQM effectively highlights anomalous regions in histograms, allowing shifters to quickly pinpoint issues that may affect data quality and subsequent physics analyses.

The performance of AutoDQM was evaluated using a dataset of 265 good and 43 bad runs, with the tool successfully identifying over 50% of the bad runs while flagging less than 15% of the good runs as anomalous. The integration of PCA and AE allows for robust anomaly detection without the need for labeled training data, making the system adaptable to various types of anomalies. The modified chi-squared metric used in AutoDQM mitigates biases related to histogram occupancy, ensuring that the anomaly detection remains effective across runs of varying lengths. Overall, AutoDQM represents a significant advancement in automated DQM, facilitating rapid identification of detector issues and enhancing the reliability of data collected for physics research.

شارك: