الهروب من الغابة: بديل شبكي عصبي نادر، قابل للتفسير، وأساسي للبيانات الجدولية
Escaping the forest: a sparse, interpretable, and foundational neural network alternative for tabular data

شارك:
المجلة: npj Artificial Intelligence، المجلد: 2، العدد: 1
DOI: https://doi.org/10.1038/s44387-025-00056-0
تاريخ النشر: 2026-01-27
المؤلف: Salvatore Raieli وآخرون
الموضوع الرئيسي: علم النسخ الجيني أحادي الخلية والمكاني

نظرة عامة

تقدم البحث sTabNet، وهو إطار عمل مبتكر للتوليد الميتا مصمم لمجموعات البيانات الجدولية المستخدمة عادة في الأبحاث البيولوجية. على الرغم من التقدم في الهياكل العصبية، لا يزال هناك عدم توافق حول النماذج التي توازن بشكل فعال بين الأداء، وقابلية التفسير، والكفاءة. يعالج sTabNet هذه الفجوة من خلال دمج مكونين رئيسيين: توليد الهيكل تلقائيًا من خلال المشي العشوائي غير المراقب، الذي يركز على الميزات Node2Vec، مما يعزز من تشتت الشبكة والتعميم مع تقليل الإفراط في التخصيص، وطبقة انتباه مخصصة تتعلم في الوقت نفسه أهمية الميزات ومعلمات النموذج، مما يوفر قابلية تفسير داخلية.

تظهر التقييمات عبر مهام بيولوجية متنوعة – مثل تصنيف RNA-Seq، وتحليل الخلايا الفردية، وتوقع البقاء – أن sTabNet يؤدي بشكل مشابه أو أفضل من النماذج الرائدة المعتمدة على الأشجار مثل XGBoost، مع الحفاظ على الكفاءة الحسابية والتوافق مع تدريب وحدة المعالجة المركزية. يظهر الإطار عمومًا قويًا عبر مجموعات البيانات داخل المجال وخارجه، مما يوفر رؤى ذات صلة بيولوجيًا ويتفوق على طرق التفسير بعد الحدث مثل SHAP من حيث الاستقرار والوضوح. هذه النتائج تضع sTabNet كأداة أساسية ومتعددة الاستخدامات للتعلم العصبي القابل للتفسير والفعال من حيث البيانات في السياقات الجدولية.

مقدمة

في هذا القسم، يقدم المؤلفون تمديدًا لـ sTabNet لمجموعات البيانات الجدولية التي تفتقر إلى المعرفة بالمجال، لا سيما في السياقات البيولوجية حيث غالبًا ما تكون تفاعلات الميزات غير معروفة. يقترحون طريقة لتقريب الاتصال بين الميزات من خلال مشي عشوائي على رسم بياني للميزات، تم بناؤه باستخدام التشابه الكوني لتعريف الحواف بين الميزات. يتم تعريف مصفوفة الجوار \( M \) بحيث \( M_{ij} = 1 \) إذا كان التشابه الكوني بين الميزات \( i \) و \( j \) يتجاوز 0.5 أو أقل من -0.5، و \( 0 \) خلاف ذلك. من خلال استخدام Node2Vec للمشي العشوائي، يهدف المؤلفون إلى التقاط كل من التفاعلات المحلية والعالمية للميزات، مما يعزز من قوة النموذج ضد الضوضاء والميزات غير ذات الصلة.

يؤكد المؤلفون فرضيتهم باستخدام مجموعات بيانات مرجعية، مقارنة أداء نموذج sTabNet الخاص بهم ضد XGBoost والانحدار اللوجستي. يقومون بإجراء دراسة إلغاء لتقييم أهمية الميزات التي تم تحديدها من خلال المشي العشوائي، مما يكشف أن إزالة هذه الميزات تؤدي إلى انخفاض أكبر في أداء النموذج مقارنة بالميزات المختارة عشوائيًا. يشير هذا إلى أن عملية المشي العشوائي تحدد بشكل فعال الاتصالات الحاسمة داخل رسم بياني الميزات، مما يدعم قابلية تطبيق الهيكل في السيناريوهات التي تفتقر إلى المعرفة السابقة بالمجال.

طرق

تم تطوير نهج sTabNet من خلال إطار عمل ميتا-توليدي يقوم ببناء هيكله بناءً على العلاقات بين الميزات. يعمل في وضعين: وضع مدفوع بالمعرفة، حيث يتم استخدام المعرفة بالمجال لإنشاء مصفوفة ثنائية \( A \) تشفر ارتباطات مجموعات الميزات، ووضع غير مراقب، يستخدم تقنيات مثل التشابه الكوني ومشي Node2vec لتشكيل رسم بياني للميزات عندما تكون المعرفة السابقة غائبة. يضمن هذا الإطار أن كل خلية عصبية مخفية تتوافق مع مجموعة ميزات معينة أو حي رسومي، مما يؤدي إلى خريطة اتصال متفرقة وقابلة للتفسير. يتم إنشاء التشتت مسبقًا باستخدام مصفوفة القناع الثنائية \( A \)، التي يتم تطبيقها في منتج هادامارد مع مصفوفة الوزن، مما يضمن أن يتم اعتبار فقط الاتصالات المحددة مسبقًا بين الميزات أثناء التدريب.

يميز تصميم sTabNet عن الشبكات الكثيفة التقليدية ونماذج التقليم بعد التدريب. من خلال توليد الهيكل تلقائيًا بناءً على إما المعرفة بالمجال أو التشابه غير المراقب بين الميزات، فإنه يسمح بتعريف مبدئي للاتصال. يعزز فرض التشتت قبل التدريب من التعميم، ويقلل من الإفراط في التخصيص، ويحسن من قابلية التفسير. بالإضافة إلى ذلك، يتم دمج طبقة انتباه مخصصة لتعلم أهمية الميزات بالتزامن مع معلمات النموذج، مما يسهل قابلية التفسير الداخلية. تلتقط كل خلية عصبية التفاعلات المحلية للميزات، مما يتماشى مع العلاقات الهيكلية النموذجية في البيانات البيولوجية. على عكس الهياكل التي تعتمد فقط على الانتباه من أجل التشتت أو طرق التفسير بعد الحدث، يدمج sTabNet الانتباه مباشرة في هيكله، مما يوفر قابلية تفسير على مستوى الميزات أثناء التدريب مع الحفاظ على الكفاءة الحسابية.

نتائج

في هذا القسم، يقدم المؤلفون نتائج بحثهم حول sTabNet، وهو هيكل شبكة عصبية مبتكر مصمم للبيانات الجدولية التي تركز على الأداء، والكفاءة، وقابلية التفسير. بناءً على الأعمال السابقة مع الشبكات العصبية المقيدة بيولوجيًا، يدمج sTabNet التشتت وآلية الانتباه لتعزيز وزن الميزات أثناء التدريب. يستخدم الهيكل مصفوفة جيرة ثنائية للتحكم في تفاعلات الميزات بناءً على المسارات البيولوجية أو من خلال طرق التعلم غير المراقب، مما يقلل من عدد المعلمات القابلة للتدريب ويمنع الإفراط في التخصيص. تشير النتائج إلى أن sTabNet يتفوق على أو يطابق أداء النماذج المعتمدة على الأشجار، مثل XGBoost، عبر مجموعات بيانات بيولوجية متنوعة بينما يوفر رؤى أوضح حول أهمية الميزات دون الحاجة إلى طرق النسبة بعد الحدث.

يظهر المؤلفون أن تصميم sTabNet يسمح بتجميع فعال للميزات، سواء من خلال المعرفة السابقة بالمجال أو من خلال الأساليب المدفوعة بالبيانات، مما يؤدي إلى تقليل كبير في عدد المعلمات وتكاليف الحساب. يتم تسليط الضوء على قابلية توسيع النموذج من خلال مقاييس مقارنة ضد زيادة أبعاد الميزات، مما يظهر كفاءته في الفضاءات عالية الأبعاد. بالإضافة إلى ذلك، تؤكد دراسة الإلغاء أن التشتت المستحث يعمل كنوع من التنظيم، مع تأثير ضئيل على الأداء عند حذف تقنيات التنظيم التقليدية. بشكل عام، يمثل sTabNet تقدمًا واعدًا في تطوير نماذج قابلة للتفسير وفعالة للبيانات الجدولية المعقدة، لا سيما في التطبيقات البيولوجية.

نقاش

يؤكد قسم النقاش في الورقة على التحديات المتعلقة بقابلية التفسير في نماذج التعلم الآلي، لا سيما في سياق أهمية الميزات. يبرز المؤلفون غياب الحقيقة الأساسية لأهمية الميزات، مما يعقد تقييم طرق التفسير. لمعالجة ذلك، استخدموا مجموعات بيانات صناعية ذات تعقيد متحكم فيه لتقييم نهجهم. تشير نتائجهم إلى أنه بينما تكافح الطرق التقليدية مثل قيم SHAP لتمييز الميزات المفيدة عن الضوضاء، يظهر نموذجهم المقترح، sTabNet، أداءً محسنًا في تحديد الميزات المهمة من خلال آلية الانتباه المدمجة. تعزز هذه الآلية من أهمية الميزات وتحافظ أيضًا على الاستقرار في تمييز الميزات المفيدة عن الضوضاء عبر مجموعات بيانات متنوعة.

علاوة على ذلك، قام المؤلفون بتقييم sTabNet على مجموعات بيانات متعددة الأومكس المعقدة في العالم الحقيقي ووجدوا أنه يتنافس مع النماذج المعروفة مثل XGBoost. لاحظوا أن آلية الانتباه في sTabNet تحدد بشكل فعال الميزات ذات الصلة بيولوجيًا، مما يتماشى مع الجينات المعروفة المرتبطة بالسرطان. كما أظهر النموذج قدرات قوية في التعميم، حيث تكيف بنجاح مع مهام التصنيف المختلفة ومجموعات البيانات. تؤكد النتائج على إمكانية sTabNet كأداة متعددة الاستخدامات لمجموعة متنوعة من التطبيقات البيولوجية، بما في ذلك تحليل البقاء وتقنية الخلايا الفردية، بينما تعالج أيضًا مخاوف قابلية التفسير التي غالبًا ما تصاحب نماذج الشبكات العصبية. بشكل عام، يدعو المؤلفون إلى sTabNet كحل قوي لمواجهة مجموعات البيانات عالية الأبعاد والضوضاء في مجالات علمية متنوعة.

Journal: npj Artificial Intelligence, Volume: 2, Issue: 1
DOI: https://doi.org/10.1038/s44387-025-00056-0
Publication Date: 2026-01-27
Author(s): Salvatore Raieli et al.
Primary Topic: Single-cell and spatial transcriptomics

Overview

The research introduces sTabNet, a novel meta-generative framework designed for tabular datasets commonly used in biomedical research. Despite advancements in neural architectures, a consensus on models that effectively balance performance, interpretability, and efficiency remains elusive. sTabNet addresses this gap by incorporating two main components: automated architecture generation through unsupervised, feature-centric Node2Vec random walks, which enhances network sparsity and generalization while reducing overfitting, and a dedicated attention layer that concurrently learns feature importance and model parameters, thereby providing intrinsic interpretability.

Evaluations across various biomedical tasks—such as RNA-Seq classification, single-cell profiling, and survival prediction—demonstrate that sTabNet performs comparably to or better than leading tree-based models like XGBoost, while maintaining computational efficiency and compatibility with CPU training. The framework shows robust generalization across both in-domain and out-of-domain datasets, yielding biologically relevant insights and outperforming post-hoc explainability methods like SHAP in terms of stability and clarity. These findings position sTabNet as a foundational and versatile tool for data-efficient, interpretable neural learning in tabular contexts.

Introduction

In this section, the authors introduce an extension of sTabNet for tabular datasets lacking domain knowledge, particularly in biological contexts where feature interactions are often unknown. They propose a method to approximate feature connectivity through a random walk on a feature graph, constructed using cosine similarity to define edges between features. The adjacency matrix \( M \) is defined such that \( M_{ij} = 1 \) if the cosine similarity between features \( i \) and \( j \) exceeds 0.5 or is less than -0.5, and \( 0 \) otherwise. By employing Node2Vec for random walks, the authors aim to capture both local and global feature interactions, enhancing the model’s robustness against noise and irrelevant features.

The authors validate their hypothesis using benchmark datasets, comparing the performance of their sTabNet model against XGBoost and logistic regression. They conduct an ablation study to assess the significance of features identified through random walks, revealing that the removal of these features leads to a more substantial decline in model performance compared to randomly selected features. This indicates that the random walk process effectively identifies critical connections within the feature graph, supporting the architecture’s applicability in scenarios devoid of prior domain knowledge.

Methods

The sTabNet approach is developed through a meta-generative framework that constructs its architecture based on the relationships among features. It operates in two modes: a knowledge-driven mode, where domain knowledge is utilized to create a binary matrix \( A \) that encodes feature-group associations, and an unsupervised mode, which employs techniques like cosine similarity and Node2vec random walks to form a feature graph when prior knowledge is absent. This framework ensures that each hidden neuron corresponds to a specific feature group or graph neighborhood, leading to a sparse and interpretable connectivity map. The sparsity is established a priori using the binary masking matrix \( A \), which is applied in a Hadamard product with the weight matrix, ensuring that only predefined feature connections are considered during training.

The design of sTabNet distinguishes it from traditional dense networks and post-training pruning models. By generating the architecture automatically based on either domain knowledge or unsupervised feature similarity, it allows for a principled definition of connectivity. The imposition of sparsity before training enhances generalization, reduces overfitting, and improves interpretability. Additionally, a dedicated attention layer is integrated to learn feature importance concurrently with model parameters, facilitating intrinsic interpretability. Each neuron captures localized feature interactions, aligning the model with the structured relationships typical in biological data. Unlike architectures that rely solely on attention for sparsity or post-hoc interpretability methods, sTabNet incorporates attention directly into its architecture, thereby providing feature-level interpretability during training while maintaining computational efficiency.

Results

In this section, the authors present the results of their research on sTabNet, a novel neural network architecture designed for tabular data that emphasizes performance, efficiency, and interpretability. Building on previous work with biologically constrained neural networks, sTabNet integrates sparsity and an attention mechanism to enhance feature weighting during training. The architecture employs a binary adjacency matrix to control feature interactions based on biological pathways or through unsupervised learning methods, thereby reducing the number of trainable parameters and preventing overfitting. The results indicate that sTabNet outperforms or matches the performance of tree-based models, such as XGBoost, across various biomedical datasets while providing clearer insights into feature importance without the need for post-hoc attribution methods.

The authors demonstrate that sTabNet’s design allows for effective grouping of features, either through prior domain knowledge or data-driven approaches, which leads to a significant reduction in parameter count and computational costs. The model’s scalability is highlighted through benchmarks against increasing feature dimensionalities, showcasing its efficiency in high-dimensional spaces. Additionally, an ablation study confirms that the induced sparsity acts as a form of regularization, with minimal impact on performance when traditional regularization techniques are omitted. Overall, sTabNet represents a promising advancement in the development of interpretable and efficient models for complex tabular data, particularly in biomedical applications.

Discussion

The discussion section of the paper emphasizes the challenges of interpretability in machine learning models, particularly in the context of feature importance. The authors highlight the absence of a ground truth for feature importance, which complicates the evaluation of interpretability methods. To address this, they employed synthetic datasets with controlled complexity to assess their approach. Their findings indicate that while traditional methods like SHAP values struggle to distinguish informative features from noise, their proposed model, sTabNet, demonstrates improved performance in identifying significant features through an integrated attention mechanism. This mechanism not only enhances feature relevance but also maintains stability in distinguishing informative from noisy features across various datasets.

Furthermore, the authors evaluated sTabNet on complex real-world multi-omics datasets and found it to be competitive with established models like XGBoost. They noted that sTabNet’s attention mechanism effectively identifies biologically relevant features, aligning with known cancer-related genes. The model also exhibited strong generalization capabilities, successfully adapting to different classification tasks and datasets. The results underscore sTabNet’s potential as a versatile tool for various biological applications, including survival analysis and single-cell technology, while also addressing the interpretability concerns that often accompany neural network models. Overall, the authors advocate for sTabNet as a robust solution for tackling high-dimensional, noisy datasets in diverse scientific fields.

شارك: