نظام بيومتري متعدد الوسائط آمن وقابل للتفسير باستخدام دمج موثوق يعتمد على الثقة للوجه وبصمة الإصبع
A secure and explainable multimodal biometric system using trust adaptive fusion for face and fingerprint

شارك:
المجلة: Scientific Reports، المجلد: 16، العدد: 1
DOI: https://doi.org/10.1038/s41598-026-43252-x
PMID: https://pubmed.ncbi.nlm.nih.gov/41857140
تاريخ النشر: 2026-03-19
المؤلف: Pavani Chitrapu وآخرون
الموضوع الرئيسي: التعرف البيومتري والأمان

نظرة عامة

تتناول ورقة البحث تطوير إطار عمل موثوق وقابل للتفسير للمصادقة البيومترية متعددة الوسائط يستفيد من التعلم العميق لتعزيز الدقة والموثوقية. تعمل البيومترية متعددة الوسائط، التي تدمج خصائص بيومترية متعددة، على تحسين أمان المصادقة ولكنها تقدم تعقيدات بسبب الحاجة إلى مستشعرات متعددة وطرق استخراج ميزات متطورة. غالبًا ما تنتج الطرق التقليدية تمثيلات محددة للنوع تفتقر إلى القوة والملاءمة من أجل دمج فعال على مستوى الميزات. بالمقابل، تولد مستخرجات الميزات المعتمدة على التعلم العميق تمثيلات أكثر موثوقية وتمييزًا، وهو أمر أساسي لأنظمة متعددة الوسائط الفعالة.

يستخدم الإطار المقترح MobileNet لاستخراج الميزات وينفذ استراتيجية دمج موثوق (TAF) تعدل مساهمة كل نوع بيومتري بناءً على جودته، مما يزيد من القوة ضد المدخلات المزعجة. يضمن النظام الخصوصية من خلال استخدام تشفير تشون-كيم-كيم-سونغ (CKKS) المتماثل، الذي يحمي البيانات البيومترية الخام أثناء المعالجة. بالإضافة إلى ذلك، يتضمن Grad-CAM لزيادة قابلية تفسير النموذج، مما يعزز الشفافية في اتخاذ القرار. تكشف التقييمات على مجموعات بيانات CASIA-FaceV5 وCASIA-FingerprintV5 عن معدل خطأ منخفض يبلغ 0.0038 لتمثيل الميزات المدمجة، مما يبرز فعالية الإطار في المصادقة البيومترية متعددة الوسائط.

مقدمة

تناقش مقدمة ورقة البحث هذه الزيادة المتزايدة في انتشار أنظمة البيومترية للمصادقة الآمنة، والتي تستخدم السمات الجسدية والسلوكية مثل بصمات الأصابع وملامح الوجه. بينما تُستخدم أنظمة البيومترية الأحادية على نطاق واسع، فإنها غالبًا ما تواجه صعوبات مع المدخلات المزعجة أو غير المكتملة. بالمقابل، تعزز أنظمة البيومترية متعددة الوسائط، التي تدمج سمات متعددة، دقة التعرف ومرونة ضد الهجمات الاحتيالية وهجمات العرض. تقتصر طرق استخراج الميزات التقليدية، مثل نمط البكسل الثنائي المحلي (LBP) وتحويل الميزات غير القابلة للتغيير (SIFT)، على حساسيتها للتغيرات البيئية.

لقد حسنت التقدمات الأخيرة في التعلم العميق (DL)، وخاصة من خلال الشبكات العصبية التلافيفية (CNNs) وآليات الانتباه، بشكل كبير القدرة على التعلم من البيانات الخام من خلال التركيز على المناطق الحرجة داخل الصور المدخلة. تؤكد الورقة على أهمية تأمين المعلومات البيومترية، حيث لا يمكن إعادة تعيينها مثل كلمات المرور. لمعالجة مخاوف الأمان، تم اقتراح مجموعة متنوعة من مخططات حماية القوالب البيومترية، بما في ذلك البيومترية القابلة للإلغاء والتشفير المتماثل (HE). يُعتبر HE جديرًا بالملاحظة بشكل خاص لقدرته على إجراء حسابات على البيانات المشفرة، مما يحافظ على الخصوصية. تبرز الورقة استخدام متغير CKKS من التشفير المتماثل الكامل (FHE) في النظام المقترح، حيث يُعتبر الأكثر كفاءة لمطابقة الميزات البيومترية الآمنة.

طرق

يستخدم النظام المقترح بنية عميل-خادم تتميز بمصادق موثوق (TA) لتسهيل المصادقة البيومترية متعددة الوسائط بشكل آمن باستخدام سمات الوجه وبصمات الأصابع. تتكون المنهجية من مرحلتين رئيسيتين: التسجيل والتحقق. خلال التسجيل، يلتقط جهاز العميل صورًا للوجه وبصمات الأصابع، مستخرجًا ميزات تمييزية من خلال نموذج MN+CA، الذي يولد أيضًا درجات ثقة محددة للنوع. تخضع هذه الميزات لعملية دمج مدركة للثقة، مستندة إلى درجات الثقة، لإنشاء متجه ميزات مدمج قوي يتم تشفيره باستخدام التشفير المتماثل الكامل (FHE) وتخزينه بشكل آمن في السحابة كقالب مرجعي. في مرحلة التحقق، تتم معالجة عينات بيومترية جديدة بطريقة مماثلة، ويتم مطابقة قالب الاستعلام الناتج ضد القالب المرجعي المشفر بواسطة TA، الذي يحدد قبول المستخدم بناءً على عتبة محددة مسبقًا.

تم إجراء تقييمات تجريبية على محطة عمل مخصصة لضمان تقييم موثوق للأداء، باستخدام معالج Intel Core i7 وبطاقة رسومات NVIDIA GeForce RTX 3060. تشير النتائج إلى أن نظام دمج الوجه وبصمة الإصبع المقترح يحقق دقة عالية، حيث تتجاوز معظم التكوينات 99% من الدقة وتظهر فصلًا ممتازًا بين الفئات، كما يتضح من قيم ROC-AUC القريبة من 1.0. أظهر تكوين الوجه + L1 أدنى معدل خطأ متساوي (EER) وأفضل أداء للتحقق، بينما أظهر تكوين الوجه + L3 فعالية أقل، مما يبرز أهمية اختيار الميزات. بشكل عام، تؤكد النتائج قدرة النظام على تحقيق توازن فعال بين الأمان وسهولة الاستخدام، مع أداء متسق عبر مقاييس التقييم المختلفة.

نتائج

تظهر نتائج تصورات Grad-CAM قابلية تفسير النظام المقترح عبر كل من أنماط الوجه وبصمات الأصابع. بالنسبة لنمط الوجه، يتركز انتباه النموذج بشكل أساسي على منطقة محيط العينين (حول العينين)، وجسر الأنف، والفم، وهي ميزات حاسمة للتعرف الشخصي. من الجدير بالذكر أن النموذج يتجاهل المعلومات غير ذات الصلة مثل الشعر وعناصر الخلفية، مما يدل على اعتماده على الخصائص الوجهية المهمة.

في نمط بصمة الإصبع، يظهر النموذج أيضًا نهجًا مركزًا، حيث يركز على الحواف والمناطق المركزية لبصمة الإصبع. يحدد بفعالية نهايات الحواف والتفرعات، وهي ميزات أساسية في التعرف على بصمات الأصابع. كما يتجاهل النموذج بنجاح الضوضاء الخلفية وعيوب الصورة، مما يبرز قدرته على إعطاء الأولوية للمعلومات ذات الصلة من أجل التعرف الدقيق.

مناقشة

تسلط قسم المناقشة في ورقة البحث الضوء على قيود أنظمة البيومترية القائمة على التعلم العميق (DL)، ولا سيما نقص الشفافية والأمان. تعزز طريقة الدمج المعتمدة على الثقة (TAF) موثوقية التعرف البيومتري من خلال وزن المدخلات ديناميكيًا بناءً على موثوقيتها، مما يؤدي إلى تحسين الأداء حتى في وجود بيانات مزعجة أو منخفضة الجودة. تشمل المساهمات الرئيسية للنظام دمج MobileNet مع الانتباه القنوي لاستخراج الميزات بشكل فعال، وتنفيذ مخطط تشفير CKKS المتماثل للحفاظ على الخصوصية لمطابقة آمنة، واستخدام Grad-CAM لزيادة قابلية تفسير النموذج.

تستعرض الورقة الأعمال ذات الصلة في أنظمة البيومترية متعددة الوسائط، مشددة على أنه بينما حسنت العديد من الأساليب الدقة، فإنها غالبًا ما تعالج جميع الأنماط بشكل متساوٍ، مما قد يؤدي إلى ثغرات في البيئات المزعجة. يعالج النظام المقترح هذه القضايا من خلال دمج آلية تعتمد على الثقة تعدل مساهمة كل نمط بناءً على موثوقيته، مما يعزز القوة في التطبيقات الواقعية. علاوة على ذلك، يوفر استخدام Grad-CAM الشفافية، مما يسمح للمستخدمين بفهم عملية اتخاذ القرار للنموذج، وبالتالي زيادة الثقة في نظام المصادقة البيومترية.

Journal: Scientific Reports, Volume: 16, Issue: 1
DOI: https://doi.org/10.1038/s41598-026-43252-x
PMID: https://pubmed.ncbi.nlm.nih.gov/41857140
Publication Date: 2026-03-19
Author(s): Pavani Chitrapu et al.
Primary Topic: Biometric Identification and Security

Overview

The research paper discusses the development of a secure and explainable multimodal biometric authentication framework that leverages deep learning to enhance accuracy and reliability. Multimodal biometrics, which integrate multiple biometric characteristics, improve authentication security but introduce complexities due to the need for multiple sensors and sophisticated feature extraction methods. Traditional methods often yield modality-specific representations that lack robustness and compatibility for effective feature-level fusion. In contrast, deep learning feature extractors generate more reliable and discriminative representations, essential for effective multimodal systems.

The proposed framework employs MobileNet for feature extraction and implements a Trust Adaptive Fusion (TAF) strategy that adjusts the contribution of each biometric modality based on its quality, thereby increasing robustness against noisy inputs. The system ensures privacy through the use of Cheon-Kim-Kim-Song (CKKS) homomorphic encryption, which protects raw biometric data during processing. Additionally, it incorporates Grad-CAM for model interpretability, enhancing transparency in decision-making. Evaluations on the CASIA-FaceV5 and CASIA-FingerprintV5 datasets reveal a low error rate of 0.0038 for the fused feature representation, underscoring the framework’s effectiveness in multimodal biometric authentication.

Introduction

The introduction of this research paper discusses the increasing prevalence of biometric systems for secure authentication, which utilize physical and behavioral traits such as fingerprints and facial features. While unimodal biometric systems are widely used, they often struggle with noisy or incomplete inputs. In contrast, multimodal biometric systems, which integrate multiple traits, enhance recognition accuracy and resilience against spoofing and presentation attacks. Traditional feature extraction methods, such as Local Binary Pattern (LBP) and Scale-Invariant Feature Transform (SIFT), are limited by their susceptibility to environmental variations.

Recent advancements in Deep Learning (DL), particularly through Convolutional Neural Networks (CNNs) and attention mechanisms, have significantly improved the ability to learn from raw data by focusing on critical areas within input images. The paper emphasizes the importance of securing biometric information, as it cannot be reset like passwords. To address security concerns, various biometric template protection schemes have been proposed, including cancellable biometrics and homomorphic encryption (HE). HE is particularly noteworthy for its ability to perform computations on encrypted data, thereby maintaining privacy. The paper highlights the use of the CKKS variant of Fully Homomorphic Encryption (FHE) in the proposed system, as it is deemed the most efficient for secure biometric feature matching.

Methods

The proposed system employs a client-server architecture featuring a Trusted Authenticator (TA) to facilitate secure multimodal biometric authentication using facial and fingerprint traits. The methodology consists of two primary stages: enrollment and verification. During enrollment, the client device captures facial and fingerprint images, extracting discriminative features through the MN+CA model, which also generates modality-specific confidence scores. These features undergo a trust-aware fusion process, informed by confidence scores, to create a robust fused feature vector that is encrypted using Fully Homomorphic Encryption (FHE) and stored securely in the cloud as a reference template. In the verification stage, new biometric samples are processed similarly, and the resulting probe template is matched against the encrypted reference template by the TA, which determines user acceptance based on a predefined threshold.

Experimental evaluations were conducted on a dedicated workstation to ensure reliable performance assessment, utilizing an Intel Core i7 processor and an NVIDIA GeForce RTX 3060 GPU. The results indicate that the proposed face-fingerprint fusion system achieves high accuracy, with most configurations exceeding 99% accuracy and demonstrating excellent class separability, as evidenced by ROC-AUC values close to 1.0. The Face + L1 configuration exhibited the lowest Equal Error Rate (EER) and optimal verification performance, while the Face + L3 configuration showed less effectiveness, highlighting the importance of feature selection. Overall, the findings affirm the system’s capability to balance security and usability effectively, with consistent performance across various evaluation metrics.

Results

The results of the Grad-CAM visualizations demonstrate the interpretability of the proposed system across both face and fingerprint modalities. For the face modality, the model’s attention is primarily focused on the periocular area (around the eyes), the nose bridge, and the mouth, which are critical features for personal identification. Notably, the model disregards irrelevant information such as hair and background elements, indicating its reliance on significant facial characteristics.

In the fingerprint modality, the model similarly exhibits a focused approach, concentrating on the ridges and central regions of the fingerprint. It effectively identifies ridge endings and bifurcations, which are essential features in fingerprint recognition. The model also successfully ignores background noise and image artifacts, further underscoring its capability to prioritize relevant information for accurate identification.

Discussion

The discussion section of the research paper highlights the limitations of existing deep learning (DL) biometric systems, particularly their lack of transparency and security. The proposed Trust-Adaptive Fusion (TAF) method enhances the reliability of biometric recognition by dynamically weighting inputs based on their trustworthiness, leading to improved performance even in the presence of noisy or low-quality data. Key contributions of the system include the integration of MobileNet with channel attention for effective feature extraction, the implementation of a privacy-preserving CKKS homomorphic encryption scheme for secure matching, and the use of Grad-CAM for model interpretability.

The paper reviews related work in multimodal biometric systems, emphasizing that while many approaches have improved accuracy, they often treat all modalities equally, which can lead to vulnerabilities in noisy environments. The proposed system addresses these issues by incorporating a trust-adaptive mechanism that adjusts the contribution of each modality based on its reliability, thus enhancing robustness in real-world applications. Furthermore, the use of Grad-CAM provides transparency, allowing users to understand the decision-making process of the model, thereby increasing trust in the biometric authentication system.

شارك: