طريقة ذكية للكشف عن البرمجيات الخبيثة تعتمد على مقاطع الذاكرة
Intelligent malware detection method based on memory segments

شارك:
المجلة: Cybersecurity، المجلد: 9، العدد: 1
DOI: https://doi.org/10.1186/s42400-025-00537-4
تاريخ النشر: 2026-01-04
المؤلف: Songcheng Yu وآخرون
الموضوع الرئيسي: تقنيات الكشف المتقدمة عن البرمجيات الخبيثة

نظرة عامة

تتناول هذه الورقة البحثية قيود تقنيات الكشف التقليدية عن البرمجيات الضارة، لا سيما في سياق تطور متغيرات البرمجيات الضارة. لتعزيز قدرات الكشف، يقترح المؤلفون خوارزمية كشف ذكية تستخدم شبكة عصبية تلافيفية أحادية البعد (1D CNN) تعالج تسلسلات بايت الذاكرة الخام، المشار إليها بمقاطع الذاكرة. تقضي هذه الطريقة الآلية الشاملة على الحاجة إلى هندسة الميزات اليدوية، مما يسمح للخوارزمية بتعلم الميزات العميقة والعلاقات الكامنة في البيانات. تظهر النتائج التجريبية الأداء المثير للإعجاب للخوارزمية، حيث تحقق دقة تبلغ 98.28%، ودقة 98.94%، واسترجاع 97.6%، ودرجة F1 تبلغ 0.9826، وقيمة AUC تبلغ 0.9972، مما يدل على فعاليتها في تحديد مقاطع الذاكرة الضارة.

في الختام، يحسن النموذج المقترح بشكل كبير على تقنيات الكشف عن البرمجيات الضارة المعتمدة على الذاكرة الموجودة من خلال أتمتة عملية التحليل وتحديد البرمجيات الضارة بدقة من خلال ميزات مقاطع الذاكرة التمثيلية. تتضمن المنهجية معالجة صور الذاكرة، واستخراج بيانات مقاطع التعليمات البرمجية، وتحويل تسلسلات البايت إلى متجهات ميزات موحدة لإدخالها في شبكة 1D CNN. بينما يظهر النموذج قدرات كشف قوية، قد يفوت بعض البرمجيات الضارة الخفية ذات البصمات الذاكرية المنخفضة. ستركز الأبحاث المستقبلية على تعزيز قوة النموذج ضد تقنيات التعتيم وتحسين استخدام الموارد لتقليل عبء الأداء، مع معالجة تحدي الكشف عن البرمجيات الضارة المخفية بشدة. تسلط هذه الدراسة الضوء على إمكانيات دمج التعلم العميق مع الطب الشرعي للذاكرة لتطوير استراتيجيات الكشف عن البرمجيات الضارة.

مقدمة

تسلط مقدمة هذه الورقة البحثية الضوء على التهديد المتزايد للجريمة الرقمية، لا سيما من خلال البرمجيات الضارة المتطورة التي تطورت جنبًا إلى جنب مع التقدم في تقنيات الإنترنت والتحول الرقمي. مع زيادة متغيرات البرمجيات الضارة سنويًا بمئات الآلاف، فإن أنظمة الكشف التقليدية التي تعتمد على التوقيعات الثابتة والقواعد المحددة مسبقًا تثبت عدم كفايتها ضد التهديدات الجديدة والمعدلة. تؤكد الورقة على التحديات التي تطرحها تقنيات مثل التعددية والبرمجيات الضارة بدون ملفات، والتي تعقد الكشف عن التهديدات غير المعروفة وتعيق التحليل الجنائي.

لمعالجة هذه التحديات، يقترح البحث طريقة كشف جديدة تستفيد من تقنيات التعلم العميق لتحليل تسلسلات البايت الخام من صفحات الذاكرة، بدلاً من الاعتماد على التوقيعات المعروفة أو ميزات سلوكية محددة. يحدد الدراسة ثلاثة جوانب رئيسية: تحويل البيانات الثنائية غير المنظمة إلى تمثيلات فعالة للتعلم الآلي، وتصميم نموذج للتعرف على أنماط البرمجيات الضارة الدقيقة، وإقامة إطار كشف شامل لتعزيز الأتمتة وقدرات التعميم. تستخدم الطريقة المقترحة شبكة عصبية تلافيفية أحادية البعد (1D CNN) لتعلم وتحديد الأنماط الضارة تلقائيًا، مما يحسن كفاءة الكشف عن البرمجيات الضارة ويسهل تصنيف عائلات البرمجيات الضارة. تشمل المساهمات الرئيسية إنشاء مجموعة بيانات صور الذاكرة، وتقديم مفهوم مقطع الذاكرة للتحليل الجنائي الذكي، وتقييم منهجي لقدرات الكشف للنموذج المقترح مقارنة بالطرق الموجودة.

الطرق

في هذه الدراسة، تم تطوير نموذج شبكة عصبية تلافيفية أحادية البعد (CNN) للكشف عن البرمجيات الضارة باستخدام ميزات مقاطع الذاكرة. تضمنت الإعدادات التجريبية آلة افتراضية VMware لجمع البيانات ومعالجتها، مع تدريب النموذج على مضيف محلي مزود بوحدة معالجة مركزية بـ 16 نواة، وذاكرة وصول عشوائي 32 جيجابايت، وGPU RTX 5080. شملت المنهجية تنفيذ عينات من البرمجيات الضارة وغير الضارة في دفعات، والتقاط صور الذاكرة، واستخراج بيانات مقاطع التعليمات البرمجية، والتي تم تحويلها بعد ذلك إلى نص سداسي عشري ومعالجتها إلى مصفوفات بايت بطول ثابت. تضمنت مجموعة البيانات حوالي 200,000 عينة غير ضارة و200,000 عينة ضارة، مما يضمن توزيعًا متوازنًا من خلال أخذ عينات طبقية للتدريب والتحقق والاختبار.

حقق النموذج مقاييس أداء مثيرة للإعجاب، بما في ذلك دقة قصوى تبلغ 98.28%، ودقة 98.94%، واسترجاع 97.6%، ودرجة F1 تبلغ 0.9826، وقيمة AUC تبلغ 0.9972 على مدار 50 دورة تدريبية. أظهرت المقارنات مع النماذج الموجودة، مثل CNN-Isolation Forest وطريقة Zhang et al.، سرعة التنبؤ المتفوقة للنموذج المقترح وقدرات الكشف الفعالة. أنشأت الدراسة عتبة تبلغ 20% لنسبة مقاطع الذاكرة الضارة المكتشفة للإشارة إلى هجوم برمجي ضار، تم التحقق منها من خلال اختبارات تجريبية متعددة، مما يكشف عن اختلافات كبيرة في مقاطع الذاكرة بين الأنظمة التي تحتوي على برمجيات ضارة والتي لا تحتوي عليها. تسلط هذه الأبحاث الضوء على فعالية استخدام مقاطع الذاكرة كميزات للكشف عن البرمجيات الضارة، مما يساهم في فهم التطبيقات العميقة للتعلم في الأمن السيبراني.

المناقشة

تسلط قسم المناقشة في الورقة الضوء على تطور الحالة الحالية للطب الشرعي للذاكرة وكشف البرمجيات الضارة، مع التركيز على قيود الطرق التقليدية التي تعتمد بشكل كبير على خبرة المحللين وقواعد الكشف المحددة مسبقًا. بينما تم تعديل الفحص القائم على الميزات الثابتة لتحليل الذاكرة، فإنه يواجه صعوبة ضد تقنيات التعتيم المتطورة للبرمجيات الضارة. في المقابل، أظهرت الأساليب المعتمدة على السلوك باستخدام التعلم الآلي وعدًا من خلال استخراج ميزات دلالية على مستوى أعلى، ومع ذلك لا تزال تواجه تحديات في هندسة الميزات وإعادة بناء المعلومات الدلالية بدقة. كما تشير الورقة إلى التطبيق المتزايد للتعلم العميق في الأمن السيبراني، لا سيما في الكشف عن البرمجيات الضارة، حيث أظهرت الشبكات العصبية التلافيفية (CNNs) أداءً متفوقًا في تحديد الأنماط المعقدة في البيانات.

تستفيد طريقة تحليل البرمجيات الضارة المقترحة من إطار جنائي يعتمد على مقاطع الذاكرة، باستخدام شبكة 1D CNN لتحليل ميزات الذاكرة المستخرجة. تشمل المنهجية عدة مراحل: اكتساب البيانات، المعالجة، تدريب النموذج، وكشف مقاطع الذاكرة الضارة. يتم تدريب النموذج على مجموعة بيانات مستمدة من مجموعة بيانات BOD-MAS، محققًا مقاييس مثيرة للإعجاب مثل دقة 98.28%، ودقة 98.94%، واسترجاع 97.6%، ودرجة F1 تبلغ 0.9826. تؤكد هذه النتائج فعالية استخدام مقاطع الذاكرة كمجموعة ميزات للكشف عن البرمجيات الضارة، مما يظهر قوة النموذج وقدرته على التعميم عبر مجموعات بيانات متنوعة. تختتم الورقة بالاعتراف بإمكانات الطريقة في تعزيز الطب الشرعي للذاكرة وكشف البرمجيات الضارة، مع تحديد مجالات البحث المستقبلية، بما في ذلك تحسين المقاومة لتقنيات التعتيم وتحسين استخدام الموارد.

القيود

تسلط القيود المفروضة على الأبحاث الحالية في الكشف عن البرمجيات الضارة الضوء على التحديات الكبيرة في الكشف عن الذاكرة في الوقت الحقيقي، لا سيما مع التقدم الأخير مثل نماذج Foundation (مثل MalConv-2، Ember-BERT). بينما تعزز هذه النماذج التقاط الميزات لملفات Portable Executable (PE) المعتمدة على القرص، فإنها غير مناسبة لخصائص تسلسل الذاكرة القصيرة، والتي عادة ما تكون حوالي 4KB، مما يؤدي إلى زيادة زمن الاستدلال والمعلمات الزائدة عند تطبيقها مباشرة. بالإضافة إلى ذلك، تواجه المحاولات لدمج ميزات مقاطع الذاكرة مع تسلسلات استدعاء API باستخدام الشبكات العصبية الرسومية مشكلات نشر عملية بسبب ارتفاع تكاليف اكتساب البيانات واستهلاك الموارد، حيث تتطلب جمع بيانات متنوعة في وقت واحد.

في المقابل، تعالج الطريقة المقترحة في هذه الورقة هذه القيود من خلال التركيز فقط على بيانات مقطع الذاكرة الفردية، مما يسهل النشر المبسط. باستخدام التعلم التبايني لاستخراج الميزات دون الحاجة إلى مجموعات بيانات تدريب مسبقة كبيرة، تتيح الطريقة تدريبًا فعالًا من البداية إلى النهاية على عينات مصنفة، مما يجعلها مناسبة لسيناريوهات العينات الصغيرة. تحدد هذه الورقة فجوة بحثية في نقص طرق الكشف من البداية إلى النهاية المصممة لتناسب دقة صفحات الذاكرة وتقدم ابتكارات في ثلاثة مجالات رئيسية: (1) دقة التحليل التي تركز على صفحات الذاكرة للكشف عن السلوكيات الضارة المحلية، (2) القضاء على هندسة الميزات اليدوية المعقدة من خلال الاستخدام المباشر لقيم البايت الخام، و(3) تطوير نموذج 1D CNN خفيف الوزن مُحسَّن لتسلسلات بايت الذاكرة، مما يعزز أداء الكشف مع الحفاظ على الكفاءة.

Journal: Cybersecurity, Volume: 9, Issue: 1
DOI: https://doi.org/10.1186/s42400-025-00537-4
Publication Date: 2026-01-04
Author(s): Songcheng Yu et al.
Primary Topic: Advanced Malware Detection Techniques

Overview

This research paper addresses the limitations of conventional malware detection techniques, particularly in the context of evolving malware variants. To enhance detection capabilities, the authors propose an intelligent detection algorithm utilizing a one-dimensional convolutional neural network (1D CNN) that processes raw memory byte sequences, referred to as memory segments. This end-to-end automated approach eliminates the need for manual feature engineering, allowing the algorithm to learn deep features and relationships inherent in the data. The experimental results demonstrate the algorithm’s impressive performance, achieving an accuracy of 98.28%, precision of 98.94%, recall of 97.6%, F1-score of 0.9826, and an AUC value of 0.9972, indicating its effectiveness in identifying malicious memory segments.

In conclusion, the proposed model significantly improves upon existing memory-based malware detection technologies by automating the analysis process and accurately identifying malware through representative memory segment features. The methodology involves preprocessing memory images, extracting code segment data, and transforming byte sequences into standardized feature vectors for input into the 1D CNN. While the model shows strong detection capabilities, it may still miss certain stealthy malware with low memory footprints. Future research will focus on enhancing the model’s robustness against obfuscation techniques and optimizing resource utilization to minimize performance overhead, while also addressing the challenge of detecting highly concealed malware. This study highlights the potential of integrating deep learning with memory forensics to advance malware detection strategies.

Introduction

The introduction of this research paper highlights the escalating threat of digital crime, particularly through sophisticated malware that has evolved alongside advancements in internet technologies and digital transformation. With malware variants increasing annually by hundreds of thousands, traditional detection systems that rely on static signatures and predefined rules are proving inadequate against new and modified threats. The paper emphasizes the challenges posed by techniques such as polymorphism and fileless malware, which complicate the detection of unknown threats and hinder forensic analysis.

To address these challenges, the research proposes a novel detection method that leverages deep learning techniques to analyze raw byte sequences from memory pages, rather than relying on known signatures or specific behavioral features. The study outlines three key aspects: transforming unstructured binary data into effective representations for machine learning, designing a model to recognize subtle malware patterns, and establishing an end-to-end detection framework to enhance automation and generalization capabilities. The proposed method utilizes a one-dimensional convolutional neural network (1D CNN) to automatically learn and identify malicious patterns, thereby improving the efficiency of malware detection and facilitating the classification of malware families. The main contributions include the creation of a memory image dataset, the introduction of the memory segment concept for intelligent forensic analysis, and a systematic evaluation of the proposed model’s detection capabilities compared to existing methods.

Methods

In this study, a 1D Convolutional Neural Network (CNN) model was developed for malware detection using memory segment features. The experimental setup involved a VMware virtual machine for dataset collection and preprocessing, with model training conducted on a local host equipped with a 16-core CPU, 32GB RAM, and an RTX 5080 GPU. The methodology included executing benign and malicious software samples in batches, capturing memory images, and extracting code segment data, which was then transformed into hexadecimal text and processed into fixed-length byte arrays. The dataset comprised approximately 200,000 benign and 200,000 malicious samples, ensuring a balanced distribution through stratified sampling for training, validation, and testing.

The model achieved impressive performance metrics, including a maximum accuracy of 98.28%, precision of 98.94%, recall of 97.6%, F1-score of 0.9826, and an AUC value of 0.9972 over 50 training epochs. Comparisons with existing models, such as the CNN-Isolation Forest and Zhang et al.’s method, demonstrated the proposed model’s superior prediction speed and effective detection capabilities. The study established a threshold of 20% for the ratio of detected malicious memory segments to indicate a malware attack, validated through multiple experimental tests, revealing significant differences in memory segments between systems with and without executed malware. This research highlights the effectiveness of utilizing memory segments as features for malware detection, contributing to the field’s understanding of deep learning applications in cybersecurity.

Discussion

The discussion section of the paper highlights the evolution and current state of memory forensics and malware detection, emphasizing the limitations of traditional methods that rely heavily on analyst experience and predefined detection rules. While static feature-based scanning has been adapted for memory analysis, it struggles against sophisticated malware obfuscation techniques. In contrast, behavior-based approaches utilizing machine learning have shown promise by extracting higher-level semantic features, yet they still face challenges in feature engineering and the accurate reconstruction of semantic information. The paper also notes the growing application of deep learning in cybersecurity, particularly in malware detection, where convolutional neural networks (CNNs) have demonstrated superior performance in identifying complex patterns in data.

The proposed malware analysis method leverages a forensic framework based on memory segments, utilizing a 1D CNN to analyze extracted memory features. The methodology includes several phases: data acquisition, processing, model training, and detection of malicious memory segments. The model is trained on a dataset derived from the BOD-MAS dataset, achieving impressive metrics such as an accuracy of 98.28%, precision of 98.94%, recall of 97.6%, and an F1-score of 0.9826. These results underscore the effectiveness of using memory segments as a feature set for malware detection, demonstrating the model’s robustness and generalization across various datasets. The paper concludes by acknowledging the method’s potential in enhancing memory forensics and malware detection, while also identifying areas for future research, including improving resistance to obfuscation techniques and optimizing resource utilization.

Limitations

The limitations of existing research in malware detection highlight significant challenges in real-time memory detection, particularly with recent advancements such as Foundation models (e.g., MalConv-2, Ember-BERT). While these models enhance feature capture for disk-based Portable Executable (PE) files, they are ill-suited for the short-sequence characteristics of memory segments, typically around 4KB, leading to increased inference latency and redundant parameters when applied directly. Additionally, attempts to fuse memory segment features with API call sequences using graph neural networks face practical deployment issues due to high data acquisition costs and resource consumption, as they require simultaneous collection of diverse data sources.

In contrast, the proposed method in this paper addresses these limitations by focusing solely on single memory segment data, facilitating streamlined deployment. Utilizing contrastive learning for feature extraction without the need for extensive pre-training datasets, the approach allows for efficient end-to-end training on labeled samples, making it suitable for small-sample scenarios. This paper identifies a research gap in the lack of end-to-end detection methods tailored to memory page granularity and presents innovations in three key areas: (1) analysis granularity focused on memory pages to detect local malicious behaviors, (2) elimination of complex manual feature engineering through direct use of raw byte values, and (3) the development of a lightweight 1D CNN model optimized for memory byte sequences, enhancing detection performance while maintaining efficiency.

شارك: