تقييم فعالية نماذج التعلم الآلي في توقع الأداء في كرة السلة: دراسة مقارنة
Evaluating the effectiveness of machine learning models for performance forecasting in basketball: a comparative study

المجلة: Knowledge and Information Systems، المجلد: 66، العدد: 7
DOI: https://doi.org/10.1007/s10115-024-02092-9
تاريخ النشر: 2024-03-24
المؤلف: George Papageorgiou وآخرون
الموضوع الرئيسي: تحليلات الرياضة والأداء

نظرة عامة

تستكشف هذه الورقة البحثية تطبيق نماذج التعلم الآلي (ML) لتوقع أداء اللاعبين في كرة السلة، مع التركيز بشكل خاص على 14 نموذجًا مختلفًا من ML تم تطبيقها على 18 إحصائية متقدمة في كرة السلة مستمدة من مجموعة من 90 لاعبًا عالي الأداء في الدوري الأمريكي للمحترفين (NBA). تقدم الدراسة مقياس تقييم جديد، وهو نسبة الخطأ المئوية المتوسطة الموزونة (WAPE)، لتقييم دقة التوقعات لهذه النماذج. تشير النتائج إلى أن النماذج المعتمدة على الأشجار، وخاصة الأشجار الإضافية (ET)، والغابات العشوائية (RF)، وشجرة القرار (DT)، تتفوق باستمرار على النماذج الأخرى، حيث حققت ET أفضل أداء بنسبة WAPE تبلغ 34.14%. كما تشير الدراسة إلى تحسن بنسبة 3.6% في متوسط نسبة الخطأ المطلق (MAPE) لمؤشرات الأداء الرئيسية (KPIs) عند استخدام نهج شامل للتوقع.

في الختام، تسلط الدراسة الضوء على التحديات وأهمية توقع الأداء في كرة السلة، مع التأكيد على فعالية استخدام نماذج ML متعددة مصممة لتناسب إحصائيات مستهدفة معينة. تكشف النتائج أن النماذج المعتمدة على الأشجار تتفوق في توقع أداء اللاعبين، بينما تظهر النماذج الخطية مثل LASSO أيضًا قدرات قوية. علاوة على ذلك، تقترح الدراسة أن استراتيجية التوقع الأكثر فعالية تتضمن توقع الإحصائيات الفردية التي تساهم في KPIs بشكل منفصل، مما يؤدي إلى تحسين دقة التوقعات بشكل عام. تؤكد النتائج على الإمكانية لتحسين نتائج التوقع في مقاييس أداء كرة السلة من خلال منهجيات توقع مبتكرة.

مقدمة

تسلط مقدمة هذه الورقة البحثية الضوء على المجال المتنامي لتحليلات الرياضة (SA)، وخاصة في كرة السلة، حيث يتم تطبيق تقنيات التعلم الآلي (ML) والتنقيب عن البيانات (DM) بشكل متزايد لتعزيز تقييم الأداء، والوقاية من الإصابات، واتخاذ القرارات الاستراتيجية. تؤكد الدراسة على ضرورة وجود بيانات صالحة، تم جمعها من خلال وسائل مختلفة مثل الكاميرات وأجهزة الاستشعار، لدفع التقدم في SA. تحدد مؤشرات الأداء الرئيسية (KPIs) مثل الكفاءة (EFF)، ونتيجة المباراة (GMSC)، وتصنيف كفاءة اللاعب (PER) كمقاييس أساسية لتقييم أداء اللاعبين، وهي ضرورية للفرق والمنظمات في عمليات التقييم الخاصة بهم.

الهدف الرئيسي من البحث هو تقييم أداء 14 نموذجًا مختلفًا من ML في توقع أداء لاعبي كرة السلة بناءً على 18 مقياسًا متقدمًا، باستخدام بيانات من 90 لاعبًا عالي الأداء على مدار ثلاثة مواسم في الدوري الأمريكي للمحترفين (2019-2022). تستخدم الدراسة نهجًا جديدًا يتضمن 381 ميزة تأخير للعبة ومنهجية توقع شاملة، تتكون من تجربتين: واحدة تتوقع نقاط الخيال (FP) والأخرى تركز على مقاييس الأداء الفردية مثل النقاط، والارتدادات، والتمريرات الحاسمة. تشير النتائج إلى أن توسيع خيارات التوقع يعزز بشكل كبير دقة توقعات KPIs، مما يوفر رؤى قيمة للمساهمين في صناعة تحليلات الرياضة. لا يقيم البحث فقط فعالية نماذج ML المختلفة ولكنه يقدم أيضًا مقياس تقييم مخصص، وهو نسبة الخطأ المئوية المتوسطة الموزونة (WAPE)، لتسهيل مقارنة دقيقة لدقة التوقع.

الطرق

توضح قسم المنهجية النهج المتبع لتحليل بيانات أداء كرة السلة المستمدة من الموقع الرسمي للدوري الأمريكي للمحترفين لمواسم 2019-20، 2020-21، و2021-22. تشمل البيانات إحصائيات تقييم وأداء متنوعة، بما في ذلك نتائج اللاعبين والفرق، التي توفر رؤى حول مقاييس الهجوم والدفاع والعمل الجماعي، بالإضافة إلى مؤشرات الأداء الرئيسية المتقدمة (KPIs). تضمنت الخطوات الأولية تنظيف البيانات وتحويلها، تلتها هندسة الميزات، حيث تم إنشاء ميزات التأخير (1، 3، 5، 7، و10 مباريات) لتسهيل توقع نماذج التعلم الآلي (ML) الانحدارية.

في مرحلة التوقع، استخدمت الدراسة 14 نموذجًا مختلفًا من ML، بما في ذلك التعزيز التكيفي (AB)، وأقرب الجيران (KNN)، وأشجار القرار (DT)، والأشجار الإضافية (ET)، وLightGBM (LGBM)، والشبكة المرنة (EN)، والغابات العشوائية (RF)، وآلات تعزيز التدرج (GBM)، والعدوانية السلبية (PA)، والانحدار البايزي (BR)، والانحدار بأقل زاوية (LARS)، والانحدار Ridge (RR)، وLASSO. تم اختبار كل نموذج ضد 18 إحصائية متقدمة لأداء كرة السلة وKPIs للاعبين الفرديين، بهدف إنشاء جدول ترتيب الأداء بناءً على نسبة الخطأ المئوية المطلقة المتوسطة (MAPE) ونسبة الخطأ المئوية المطلقة الموزونة (WAPE). ركزت التجربة النهائية على حساب KPIs بناءً على نقاط الخيال (FP)، مع دمج إحصائيات اللاعبين المختلفة لتقييم الأداء العام. يتم تلخيص سير العمل للمنهجية، من جمع البيانات إلى التوقع والتقييم، بصريًا في شكل مصاحب.

النتائج

في هذا البحث، طور المؤلفون نماذج توقع أداء لاعبي كرة السلة (BPPF) بدقة عالية لتعزيز فهم اتجاهات أداء اللاعبين باستخدام تقنيات التعلم الآلي (ML). شملت الدراسة تحليل أداء 14 نموذجًا مختلفًا من ML عبر 90 دراسة حالة من لاعبي الدوري الأمريكي للمحترفين عالي الأداء، باستخدام مجموعة بيانات مقسمة إلى 70% تدريب، 20% اختبار، و10% تحقق. كانت مؤشرات الأداء الرئيسية (KPIs) المستهدفة للتوقع تشمل مقاييس مثل نسبة التمريرات الحاسمة إلى الأخطاء (AST/TO)، والكتل (BLK)، والكفاءة (EFF)، والنقاط المسجلة (PTS)، من بين أمور أخرى. كانت تقييمات أداء النموذج تعتمد بشكل أساسي على نسبة الخطأ المئوية المطلقة المتوسطة (MAPE).

أشارت النتائج إلى أن ستة فقط من بين 14 نموذجًا اعتُبرت فعالة في توقع إحصائيات كرة السلة الفردية المتقدمة. برز نموذج الأشجار الإضافية (ET) كأفضل أداء، محققًا أدنى MAPE لعدة مقاييس، بما في ذلك نسبة AST وتقدير تأثير اللاعب (PIE). تلاه نموذج الغابات العشوائية (RF) عن كثب، متفوقًا في توقع التمريرات الحاسمة والنقاط المسجلة. بالإضافة إلى ذلك، أظهر نموذج شجرة القرار (DT) أداءً قويًا في الارتدادات والسرقات. كما أظهرت نماذج أخرى، مثل LARS وLASSO، قدرات تنبؤية ملحوظة لمقاييس معينة. بشكل عام، تؤكد النتائج على إمكانيات نماذج ML في توقع أداء لاعبي كرة السلة بدقة، مما يوفر رؤى قيمة للمدربين والمحللين في الرياضة.

المناقشة

تسلط قسم المناقشة في الورقة البحثية الضوء على تطور وأهمية التحليل الإحصائي (SA) والتعلم الآلي (ML) في توقع أداء لاعبي كرة السلة. تؤكد على أن الفرق المحترفة تستفيد من SA لتعزيز أداء اللاعبين وتحديد المشكلات التي قد لا تكون واضحة للمدربين. لقد مكنت التقدمات الأخيرة في تقنيات ML والتنقيب عن البيانات (DM) الباحثين من اقتراح طرق مبتكرة لتوقع الأداء، مثل استخدام البيانات الوظيفية النادرة ونظرية الرسوم البيانية، التي أظهرت دقة محسنة مقارنة بالطرق التقليدية. من الجدير بالذكر أن الدراسات قد حددت مؤشرات الأداء الرئيسية (KPIs) التي ترتبط بالفوز، مثل كفاءة الرماية ومقاييس الدفاع، بينما تستكشف أيضًا العلاقة بين سمات شخصية اللاعب والأداء.

تستعرض الورقة منهجيات مختلفة تم استخدامها في الأبحاث السابقة، بما في ذلك استخدام إحصائيات نتائج الصناديق المتقدمة والنماذج التنبؤية لتوقع نتائج مثل MVP وترتيب اللاعبين الدفاعيين. كما تناقش أهمية هندسة الميزات والاختيار في تطوير نماذج تنبؤية قوية، مع التركيز على ميزات تأخير اللعبة وإزالة التعدد الخطي. تهدف الدراسة إلى تقييم وترتيب نماذج ML المختلفة بناءً على فعاليتها في توقع مقاييس أداء كرة السلة، مما يساهم في عمليات اتخاذ القرار الاستراتيجية داخل الفرق والمنظمات. من خلال تحليل مجموعة بيانات شاملة من مواسم الدوري الأمريكي للمحترفين الأخيرة، تسعى الدراسة إلى إنشاء إطار لتحسين تقييمات أداء اللاعبين من خلال تحليلات تنبؤية متقدمة.

Journal: Knowledge and Information Systems, Volume: 66, Issue: 7
DOI: https://doi.org/10.1007/s10115-024-02092-9
Publication Date: 2024-03-24
Author(s): George Papageorgiou et al.
Primary Topic: Sports Analytics and Performance

Overview

This research paper investigates the application of machine learning (ML) models for forecasting player performance in basketball, specifically focusing on 14 different ML models applied to 18 advanced basketball statistics derived from a pool of 90 high-performance NBA players. The study introduces a novel evaluation metric, the weighted average percentage error (WAPE), to benchmark the forecasting accuracy of these models. Results indicate that Tree-based models, particularly Extra Trees (ET), Random Forest (RF), and Decision Tree (DT), consistently outperform other models, with ET achieving the best performance at a WAPE of 34.14%. The study also reports a 3.6% improvement in mean absolute percentage error (MAPE) for key performance indicators (KPIs) when employing a comprehensive forecasting approach.

In conclusion, the research highlights the challenges and significance of performance forecasting in basketball, emphasizing the effectiveness of using multiple ML models tailored to specific target statistics. The findings reveal that Tree-based models are superior for predicting player performance, while linear models like LASSO also demonstrate strong capabilities. Furthermore, the study suggests that a more effective forecasting strategy involves predicting individual statistics that contribute to KPIs separately, leading to improved overall forecasting accuracy. The results underscore the potential for enhanced prediction outcomes in basketball performance metrics through innovative forecasting methodologies.

Introduction

The introduction of this research paper highlights the burgeoning field of sports analytics (SA), particularly in basketball, where machine learning (ML) and data mining (DM) techniques are increasingly applied to enhance performance evaluation, injury prevention, and strategic decision-making. The study emphasizes the necessity of valid data, collected through various means such as cameras and sensors, to drive advancements in SA. It identifies key performance indicators (KPIs) like efficiency (EFF), game score (GMSC), and player efficiency rating (PER) as essential metrics for assessing player performance, which are crucial for teams and organizations in their evaluation processes.

The primary objective of the research is to assess the performance of 14 different ML models in predicting basketball player performance based on 18 advanced metrics, utilizing data from 90 high-performance players over three NBA seasons (2019-2022). The study employs a novel approach that includes 381 game-lag features and a comprehensive forecasting methodology, which consists of two experiments: one predicting Fantasy Points (FP) and the other focusing on individual performance metrics such as points, rebounds, and assists. The findings indicate that expanding forecasting options significantly enhances the accuracy of KPI predictions, thereby providing valuable insights for stakeholders in the sports analytics industry. The research not only evaluates the effectiveness of various ML models but also introduces a bespoke evaluation metric, the weighted average percentage error (WAPE), to facilitate a detailed comparison of prediction accuracy.

Methods

The methodology section details the approach taken to analyze basketball performance data sourced from the official NBA website for the seasons 2019-20, 2020-21, and 2021-22. The data encompasses various evaluation and performance statistics, including Player’s and Team’s Box Scores, which provide insights into offensive, defensive, and teamwork metrics, as well as advanced Key Performance Indicators (KPIs). The initial steps involved data cleansing and transformation, followed by feature engineering, where lag features (1, 3, 5, 7, and 10 games) were created to facilitate regression machine learning (ML) model forecasting.

In the forecasting phase, the study employed 14 different ML models, including Adaptive Boosting (AB), K-Nearest Neighbors (KNN), Decision Trees (DT), Extra Trees (ET), LightGBM (LGBM), Elastic Net (EN), Random Forest (RF), Gradient Boosting Machines (GBM), Passive Aggressive (PA), Bayesian Regression (BR), Least Angle Regression (LARS), Ridge Regression (RR), and LASSO. Each model was tested against 18 advanced basketball performance statistics and KPIs for individual players, with the aim of generating a performance ranking table based on Mean Absolute Percentage Error (MAPE) and Weighted Absolute Percentage Error (WAPE) metrics. The final experiment focused on calculating KPIs based on Fantasy Points (FP), integrating various player statistics to evaluate overall performance. The methodology’s workflow, from data collection to forecasting and evaluation, is visually summarized in an accompanying figure.

Results

In this research, the authors developed highly accurate Basketball Player Performance Forecasting (BPPF) models to enhance the understanding of player performance trends using machine learning (ML) techniques. The study involved analyzing the performance of 14 different ML models across 90 case studies of high-performance NBA players, utilizing a dataset split into 70% training, 20% testing, and 10% validation. The key performance indicators (KPIs) targeted for prediction included metrics such as assists-to-turnovers (AST/TO), blocks (BLK), efficiency (EFF), and points scored (PTS), among others. The evaluation of model performance was primarily based on the Mean Absolute Percentage Error (MAPE).

The results indicated that only six of the 14 models were deemed effective for predicting advanced individual basketball statistics. The Extra Trees (ET) Regressor emerged as the top performer, achieving the lowest MAPE for several metrics, including AST ratio and Player Impact Estimate (PIE). The Random Forest (RF) Regressor followed closely, excelling in predicting assists and points scored. Additionally, the Decision Tree (DT) Regressor showed strong performance in rebounding and steals. Other models, such as LARS and LASSO, also demonstrated notable predictive capabilities for specific metrics. Overall, the findings underscore the potential of ML models in accurately forecasting basketball player performance, providing valuable insights for coaches and analysts in the sport.

Discussion

The discussion section of the research paper highlights the evolution and significance of statistical analysis (SA) and machine learning (ML) in predicting basketball players’ performance. It emphasizes that professional teams leverage SA to enhance player performance and identify issues that may not be apparent to coaches. Recent advancements in ML and data mining (DM) techniques have enabled researchers to propose innovative methods for performance prediction, such as using sparse functional data and graph theory, which have shown improved accuracy over traditional methods. Notably, studies have identified key performance indicators (KPIs) that correlate with winning, such as shooting efficiency and defensive metrics, while also exploring the relationship between player personality traits and performance.

The paper outlines various methodologies employed in previous research, including the use of advanced box score statistics and predictive models to forecast outcomes like MVP and defensive player rankings. It also discusses the importance of feature engineering and selection in developing robust predictive models, with a focus on game-lag features and the elimination of multicollinearity. The study aims to evaluate and rank different ML models based on their effectiveness in predicting basketball performance metrics, ultimately contributing to the strategic decision-making processes within teams and organizations. By analyzing a comprehensive dataset from recent NBA seasons, the research seeks to establish a framework for improving player performance assessments through advanced predictive analytics.