التعامل مع المتغيرات المستمرة ونمذجة العلاقات غير الخطية في بيانات الرعاية الصحية: دليل عملي
Dealing with continuous variables and modelling non-linear associations in healthcare data: practical guide

شارك:
المجلة: BMJ، المجلد: 390
DOI: https://doi.org/10.1136/bmj-2024-082440
PMID: https://pubmed.ncbi.nlm.nih.gov/40670054
تاريخ النشر: 2025-07-16
المؤلف: Pedro Lopez-Ayala وآخرون
الموضوع الرئيسي: دراسات ضغط الدم وارتفاع ضغط الدم

نظرة عامة

في أبحاث الرعاية الصحية، يعد تحليل المتغيرات المستمرة مثل العمر وضغط الدم ومستويات المؤشرات الحيوية أمرًا حاسمًا. ومع ذلك، غالبًا ما يقوم العديد من الباحثين بتصنيف هذه المتغيرات أو يفترضون بشكل غير صحيح وجود علاقة خطية عند نمذجة ارتباطها بالمتغيرات التابعة. يمكن أن تؤدي هذه الممارسات إلى استنتاجات مضللة، حيث أن العلاقات غير الخطية شائعة في البيانات الواقعية. لمعالجة ذلك، يُوصى باستخدام تقنيات النمذجة المرنة، مثل السبل متعددة الحدود والسبل الكسرية، لالتقاط هذه الديناميكيات غير الخطية بدقة.

تؤكد الورقة على أن سوء إدارة المتغيرات المستمرة الشائعة في الأدبيات الطبية – وبالتحديد افتراضات الخطية والتصنيف – يجب تجنبها. بدلاً من ذلك، من الضروري دمج عدم الخطية في عملية النمذجة. يبرز المؤلفون فائدة الطرق المعروفة مثل السبل المكعبة المقيدة والسبل الكسرية، والتي تكون سهلة الاستخدام وتعزز مرونة النمذجة. لمستخدمي Stata، تسهل الأوامر مثل `{makespline}` و `{mkspline}` ملاءمة B splines وrestricted cubic splines، بينما يسمح الأمر `{fp}` بنمذجة الانحدار باستخدام السبل الكسرية. يوفر المستودع المرافق أمثلة عملية وأوامر مكتوبة من قبل المستخدمين لمساعدة الباحثين في تطبيق هذه التقنيات بشكل فعال.

الطرق

في هذا القسم، يحدد المؤلفون ثلاث طرق رئيسية لنمذجة العلاقة بين متغير مستقل مستمر ومتغير تابع. تتضمن الطريقة الأولى افتراض وجود علاقة خطية، والتي يمكن تمثيلها رياضيًا كالتالي $Y = \beta_0 + \beta_1 X + \epsilon$، حيث $Y$ هو المتغير التابع، و$X$ هو المتغير المستقل، و$\epsilon$ هو مصطلح الخطأ. تتضمن الطريقة الثانية تصنيف المتغير المستقل إلى مجموعات منفصلة، مما يمكن أن يبسط التحليل ولكنه قد يؤدي إلى فقدان المعلومات وتقليل القوة الإحصائية. أخيرًا، يناقش المؤلفون استخدام تقنيات النمذجة المرنة لاستيعاب العلاقات غير الخطية، مما يسمح بفهم أكثر دقة للبيانات. لكل طريقة قيودها الخاصة، والتي يؤكد المؤلفون أنه يجب أخذها بعين الاعتبار عند اختيار استراتيجية النمذجة المناسبة.

المناقشة

في هذا القسم، يؤكد المؤلفون على الأهمية الحاسمة للتعامل بشكل مناسب مع المتغيرات المستمرة في أبحاث الرعاية الصحية، خاصة في نمذجة الانحدار. يجادلون بأن الممارسات الشائعة، مثل تصنيف المتغيرات المستمرة أو افتراض العلاقات الخطية، يمكن أن تؤدي إلى فقدان كبير للمعلومات وسوء تفسير النتائج. توضح الورقة هذه القضايا باستخدام دراسة حالة تتعلق بالقدرة التشخيصية لمستويات الجلوكوز في السائل النخاعي (CSF) في التمييز بين التهاب السحايا البكتيري الحاد والفيروسية. يظهر المؤلفون أن استخدام تقنيات النمذجة المرنة، مثل السبل المكعبة المقيدة، يمكن أن يلتقط بشكل أفضل العلاقات غير الخطية، والتي تكون غالبًا موجودة في البيانات الواقعية.

تسلط المناقشة الضوء على أن العلاقات الخطية نادرًا ما تُلاحظ في الطبيعة، وأن فرض مثل هذه الافتراضات يمكن أن يقلل من أداء النموذج والقوة الإحصائية، خاصة بالنسبة للمتنبئين المهمين مثل العمر. يقدم المؤلفون أدلة من مراجعات منهجية تشير إلى أن العديد من الدراسات تتعامل بشكل غير كافٍ مع المتغيرات المستمرة، وغالبًا ما تلجأ إلى التصنيف، مما يمكن أن يرفع معدلات الخطأ من النوع الأول ويخفي الارتباطات الحقيقية. يدعون إلى استخدام أساليب النمذجة غير الخطية، مثل السبل الكسرية والسبل، لتعزيز دقة وقابلية تفسير النماذج التنبؤية. علاوة على ذلك، يقدم المؤلفون توصيات عملية للباحثين، داعين إياهم إلى إعطاء الأولوية لاستراتيجيات النمذجة المرنة بناءً على أهمية المتغيرات ودرجات الحرية المتاحة، مما يحسن من قوة نتائج أبحاث الرعاية الصحية.

القيود

يناقش قسم القيود الاختلافات والقيود المتعلقة بالسبل الكسرية والسبل في نمذجة العلاقات غير الخطية. بينما تعتبر كلتا الطريقتين فعالتين، تستخدم السبل وظائف متعددة الحدود المحلية المعرفة بين العقد، مما يمكن أن يؤدي إلى أشكال وظيفية معقدة وعدد كبير من المعامل، خاصة عند وجود عدة متغيرات مستمرة. قد تعيق هذه التعقيدات قابلية التفسير وتثني المستخدمين عن استخدام السبل، خاصة عندما توفر البدائل الأبسط، مثل التصنيف، معاملات أكثر وضوحًا. في المقابل، تقدم السبل الكسرية نهجًا أكثر سهولة بسبب تعبيراتها وتفسيراتها الأبسط.

على الرغم من هذه الاختلافات، فإن كلا التقنيتين تنتجان نتائج مماثلة في العديد من السياقات، كما هو ملاحظ في الدراسات السابقة. يبرز تعليق من Sauerbrei وآخرين عدم وجود أدلة حاسمة تفضل طريقة على أخرى، مما يترك اختيار تقنية النمذجة لتفضيل المحلل والأهداف المحددة للنموذج. بالإضافة إلى ذلك، قد تواجه كلتا الطريقتين صعوبة في التمديد، على الرغم من أن السبل المكعبة المقيدة يمكن أن تؤدي بشكل أفضل في سيناريوهات معينة. في النهاية، إذا تم إثبات علاقة موثوقة بين المتغيرات المستقلة والتابعة من خلال الأبحاث السابقة أو الأطر النظرية، فقد يكون النموذج البارامتري الكامل مفضلًا، مما يتجنب التعقيدات المرتبطة بالسبل والسبل الكسرية.

Journal: BMJ, Volume: 390
DOI: https://doi.org/10.1136/bmj-2024-082440
PMID: https://pubmed.ncbi.nlm.nih.gov/40670054
Publication Date: 2025-07-16
Author(s): Pedro Lopez-Ayala et al.
Primary Topic: Blood Pressure and Hypertension Studies

Overview

In healthcare research, the analysis of continuous variables such as age, blood pressure, and biomarker levels is critical. However, many researchers often categorize these variables or incorrectly assume a linear relationship when modeling their association with dependent variables. Such practices can lead to misleading conclusions, as non-linear relationships are prevalent in real-world data. To address this, the use of flexible modeling techniques, such as splines and fractional polynomials, is recommended to accurately capture these non-linear dynamics.

The paper emphasizes that the common mismanagement of continuous variables in medical literature—specifically the assumptions of linearity and categorization—should be avoided. Instead, incorporating nonlinearity into the modeling process is essential. The authors highlight the utility of established methods like restricted cubic splines and fractional polynomials, which are user-friendly and enhance modeling flexibility. For Stata users, commands such as `{makespline}` and `{mkspline}` facilitate the fitting of B splines and restricted cubic splines, while the `{fp}` command allows for regression modeling with fractional polynomials. The accompanying repository provides practical examples and user-written commands to assist researchers in applying these techniques effectively.

Methods

In this section, the authors outline three primary methods for modeling the relationship between a continuous independent variable and a dependent variable. The first approach involves assuming a linear relationship, which can be represented mathematically as $Y = \beta_0 + \beta_1 X + \epsilon$, where $Y$ is the dependent variable, $X$ is the independent variable, and $\epsilon$ is the error term. The second method entails categorizing the independent variable into discrete groups, which can simplify analysis but may lead to loss of information and reduced statistical power. Lastly, the authors discuss the use of flexible modeling techniques to accommodate non-linear relationships, allowing for a more nuanced understanding of the data. Each method has its own limitations, which the authors emphasize should be carefully considered when selecting an appropriate modeling strategy.

Discussion

In this section, the authors emphasize the critical importance of appropriately handling continuous variables in healthcare research, particularly in regression modeling. They argue that common practices, such as categorizing continuous variables or assuming linear relationships, can lead to significant loss of information and misinterpretation of results. The paper illustrates these issues using a case study involving the diagnostic ability of cerebrospinal fluid (CSF) glucose levels in differentiating between acute bacterial and viral meningitis. The authors demonstrate that employing flexible modeling techniques, such as restricted cubic splines, can better capture non-linear relationships, which are often present in real-world data.

The discussion highlights that linear relationships are rarely observed in nature, and forcing such assumptions can diminish model performance and statistical power, particularly for important predictors like age. The authors provide evidence from systematic reviews indicating that many studies inadequately handle continuous variables, often resorting to categorization, which can inflate type I error rates and obscure true associations. They advocate for the use of non-linear modeling approaches, such as fractional polynomials and splines, to enhance the accuracy and interpretability of predictive models. Furthermore, the authors offer practical recommendations for researchers, urging them to prioritize flexible modeling strategies based on the importance of the variables and the available degrees of freedom, thereby improving the robustness of healthcare research findings.

Limitations

The section on limitations discusses the differences and constraints of fractional polynomials and splines in modeling non-linear relationships. While both methods are effective, splines utilize local piecewise polynomial functions defined between knots, which can lead to complex functional forms and a high number of coefficients, particularly when multiple continuous variables are involved. This complexity may hinder interpretability and deter users from employing splines, especially when simpler alternatives, like categorization, provide more straightforward coefficients. In contrast, fractional polynomials offer a more accessible approach due to their simpler expressions and interpretations.

Despite these differences, both techniques yield similar results in many contexts, as noted in previous studies. A commentary by Sauerbrei et al. emphasizes the lack of conclusive evidence to favor one method over the other, leaving the choice of modeling technique to the analyst’s preference and the specific objectives of the model. Additionally, both methods may struggle with extrapolation, although restricted cubic splines can perform better in certain scenarios. Ultimately, if a reliable relationship between the independent and dependent variables is established through prior research or theoretical frameworks, a fully parametric model may be preferable, avoiding the complexities associated with splines and polynomials.

شارك: