DOI: https://doi.org/10.3389/frai.2025.1665874
PMID: https://pubmed.ncbi.nlm.nih.gov/41561189
تاريخ النشر: 2026-01-05
المؤلف: Jorge Vila-Tomás وآخرون
الموضوع الرئيسي: التعرف على الوجه والإدراك
نظرة عامة
في هذا القسم، يجادل المؤلفون بأن نمذجة الدماغ البصري باستخدام الشبكات العصبية العميقة (ANNs) لا تزال تحديًا غير محسوم، خاصة في مجال الرؤية منخفضة المستوى. يبرزون أسئلة حاسمة تتعلق بتقييم هذه النماذج، مثل النقاط المرجعية المناسبة داخل ANNs للتقييم السلوكي، وطبيعة آليات القراءة، وما إذا كان ينبغي اعتبارها جزءًا لا يتجزأ من نموذج الدماغ. علاوة على ذلك، يناقشون ضرورة الاختبارات ذات الصلة البيولوجية لنماذج التعلم العميق الموجهة لمهام الرؤية، مؤكدين على أهمية توافق الاختبارات الاصطناعية مع النماذج التجريبية البشرية.
لمعالجة هذه القضايا، يقدم المؤلفون مجموعة بيانات جديدة منخفضة المستوى تتضمن الخصائص المكانية اللونية ذات الصلة بالاختناق التكيفي لمسار الشبكية-V1، والذي لا يتم تمثيله في قواعد البيانات الحالية مثل BrainScore. يقترحون هذه المجموعة لتقييمات نوعية وكمية لنماذج الرؤية. يوضح المؤلفون منهجيتهم من خلال تقييم ثلاثة نماذج حديثة ذات هياكل مشابهة: نموذج بارامترى تم تحسينه من خلال أقصى تمايز، ونموذج غير بارامتري (PerceptNet) متوافق مع تشوهات الصور الذاتية البشرية، ونموذج يركز على التقسيم يشارك مشفر PerceptNet. تشير نتائجهم إلى أن النموذج البارامتري يظهر سلوكًا يتماشى بشكل وثيق مع معالجة الرؤية البشرية، كما تم تقييمه مقابل عشرة خصائص بصرية نفسية/فسيولوجية مختارة.
مقدمة
ت outlines المقدمة أهمية برنامج التقييم المقترح المستند إلى علم النفس البصري منخفض المستوى، والذي يهدف إلى تكملة الاقتراحات عالية المستوى التي تم مناقشتها في Bowers et al. (2023) و Biscione et al. (2024). على الرغم من حداثة هذا النهج داخل مجتمعات الذكاء الاصطناعي وتعلم الآلة، تشير التقييمات المستقلة الأخيرة من قبل باحثين مثل Cai et al. (2025) و Hammou et al. (2025) إلى اهتمام متزايد بأساليب مماثلة. يؤكد المؤلفون أن برنامج التقييم الخاص بهم يتجاوز المعايير التقليدية للسطوع، واللون، وتغطية التباين، مؤكدين على فهم نوعي لمنحنيات الاستجابة جنبًا إلى جنب مع المقارنات الكمية.
علاوة على ذلك، تسلط المقدمة الضوء على عمل مختبر رافال مانتيك، الذي يتماشى مع التركيز على علم النفس البصري منخفض المستوى ولكنه يختلف في تأكيده على التحليل الكمي. تم تصميم إطار التقييم المقترح لتسهيل تحديد الأخطاء المفاهيمية في نماذج التعلم العميق، مما يمكّن الباحثين من إجراء تعديلات مستنيرة على هياكل النماذج بناءً على المحفزات البصرية المختارة وخصائصها النفسية الفيزيائية. يهدف هذا النهج إلى تعزيز قابلية تفسير وفعالية أنظمة الذكاء الاصطناعي من خلال تعزيز فهم أعمق لقدراتها في معالجة الرؤية.
طرق
في هذا القسم، يتناول المؤلفون التعقيدات المتعلقة بتقييم النماذج الاصطناعية للرؤية، خاصة الاختيار بين الفسيولوجيا الاصطناعية وعلم النفس الفيزيائي للإعدادات التجريبية. يبرزون أهمية اختيار الطرق المناسبة لتصور سلوك النموذج، مشيرين إلى تقنيات راسخة من علم الأعصاب البصري، مثل تحليل المجال الاستقبالي والارتباط العكسي. يقترح المؤلفون منهجية تتضمن توليد المحفزات – صور رقمية مضبوطة بشكل جيد تعكس الخصائص الرئيسية للرؤية البشرية منخفضة المستوى – واستخدامها لتقييم استجابات النموذج. يمكن إجراء التقييم من خلال مقارنات نوعية لمنحنيات الاستجابة أو مقاييس كمية مثل ارتباط بيرسون وارتباط كيندال، اعتمادًا على توفر بيانات الحقيقة الأساسية.
تؤكد المنهجية المقترحة على أهمية اختيار الطبقة الصحيحة من أجل قراءة مخرجات النموذج، حيث قد تتوافق طبقات مختلفة مع مراحل معالجة متميزة في النظام البصري البشري. يوضح المؤلفون نهجهم باستخدام ثلاث شبكات توضيحية: نموذج BioMultiLayer، وPerceptNet، وBio U-Net، كل منها مصمم لالتقاط جوانب مختلفة من معالجة الرؤية. يوضحون إعداد التجربة، بما في ذلك استخدام صور 128 × 128 واختبارات محددة لقياس استجابات النموذج، مما يمهد الطريق للتحليلات اللاحقة لأداء هذه النماذج مقابل الظواهر البصرية المعروفة.
نتائج
في قسم “النتائج”، يقدم المؤلفون تحليلات نوعية وكمية لتوافق النماذج مع السلوك البصري البشري. تشير التقييمات النوعية، الملخصة في الجدول 2، إلى أن النموذج البارامتري، BioMultiLayer، يظهر توافقًا متفوقًا مع الاستجابات البشرية، خاصة في المناطق الخطية، بسبب دمجه للحساسيات وآليات المركز المحيط. من الجدير بالذكر أن محاكاة سلوك تمرير النطاق في سياق تجربة بلاكمور وكامبل تبرز فعالية النموذج في إعادة إنتاج استجابات شبيهة بالبشر. يتم تقييم التوافق من خلال ارتباط بيرسون وارتباط كيندال، مما يكشف أنه بينما يتفوق BioMultiLayer على PerceptNet غير البارامتري وBio-U-Net، لا تزال هناك مجالات، خاصة في تفاعل الميزات، تتطلب مزيدًا من التحسين.
تستخدم التحليل الكمي، الموضح في الشكل 13، مقاييس توافق الجذر التربيعي للخطأ (RMSE) وارتباط لتقديم مقارنة مفصلة لأداء النموذج عبر خصائص مختلفة. يحقق BioMultiLayer درجة توافق متوسطة تبلغ 0.68 ± 0.14، أعلى بكثير من النماذج الأخرى، التي سجلت 0.26 ± 0.26 لـ PerceptNet و0.20 ± 0.34 لـ Bio-U-Net. تؤكد الاختبارات الإحصائية توافق BioMultiLayer المتفوق مع السلوك البشري، بينما لا تظهر النماذج الأخرى أي اختلافات ذات دلالة بينها. يحذر المؤلفون من أنه بينما تقدم المقاييس الكمية رؤى قيمة، قد تتجاهل الفروق النوعية، مما يبرز أهمية التقييمات البصرية جنبًا إلى جنب مع التقييمات العددية. لا تساعد المنهجية المقترحة فقط في فهم توافق النموذج مع البشر، بل تقترح أيضًا طرقًا لتعزيز أنظمة الرؤية الاصطناعية من خلال تقليد خصائص معالجة الرؤية الشبيهة بالبشر.
مناقشة
في قسم المناقشة من الورقة، يقوم المؤلفون بتقييم نقدي للمطالبات التي قدمها باحثون مختلفون بشأن فعالية نماذج التعلم العميق في تقليد معالجة الرؤية البشرية. يؤكدون على أهمية الشك في البحث العلمي، مشيرين إلى شخصيات تاريخية مثل توماسو بوجيو وهوراس بارلو، الذين تساءلوا عن الأطر الأساسية لعلم الرؤية. يقترح المؤلفون أن فهم الدماغ البصري من خلال تقنيات الذكاء الاصطناعي أكثر تعقيدًا مما كان يُعتقد سابقًا، داعين إلى طرق اختبار أكثر صرامة. يقدمون اختبارًا شبيهًا باختبار تورينغ يعتمد على عشرة خصائص للرؤية البشرية منخفضة المستوى، يُطلق عليه “العشر”، لتقييم أداء النماذج الاصطناعية مقابل السلوك البصري البشري.
يحدد المؤلفون هيكل ورقتهم، مشيرين إلى أنهم سيقدمون طريقة تقييم بسيطة باستخدام المحفزات البصرية التي تعكس مبادئ علم الرؤية الكلاسيكية. يخططون لتوضيح هذه الطريقة من خلال تقييم ثلاثة نماذج مختلفة، مع تسليط الضوء على ضرورة ضمان أن هذه النماذج تظهر خصائص تكيفية شبيهة بالبشر. تثير المناقشة أيضًا قضايا أوسع تتعلق بالأهداف الحاسوبية للأنظمة البصرية والتحديات المتمثلة في تفسير نماذج التعلم العميق على أنها قابلة للتطبيق بيولوجيًا. يختتم المؤلفون بالتأكيد على الحاجة إلى تحسين تقنيات التقييم لفهم مدى قرب النماذج الاصطناعية من معالجة الرؤية البشرية، مما يسهم في التقدم في كل من الذكاء الاصطناعي وعلم الرؤية.
DOI: https://doi.org/10.3389/frai.2025.1665874
PMID: https://pubmed.ncbi.nlm.nih.gov/41561189
Publication Date: 2026-01-05
Author(s): Jorge Vila-Tomás et al.
Primary Topic: Face Recognition and Perception
Overview
In this section, the authors contend that the modeling of the visual brain using deep neural networks (ANNs) remains an unresolved challenge, particularly in the realm of low-level vision. They highlight critical questions regarding the evaluation of these models, such as the appropriate points of reference within ANNs for behavioral assessment, the nature of the read-out mechanisms, and whether these should be considered integral to the brain model. Furthermore, they discuss the necessity of biologically relevant tests for deep learning models aimed at vision tasks, emphasizing the importance of aligning artificial tests with human experimental paradigms.
To address these issues, the authors introduce a novel low-level dataset that encapsulates the spatio-chromatic properties relevant to the adaptive bottleneck of the retina-V1 pathway, which is not represented in existing databases like BrainScore. They propose this dataset for both qualitative and quantitative evaluations of vision models. The authors illustrate their methodology by assessing three recent models with similar architectures: a parametric model optimized through Maximum Differentiation, a non-parametric model (PerceptNet) aligned with human subjective image distortions, and a segmentation-focused model sharing the PerceptNet encoder. Their findings indicate that the parametric model exhibits behavior most closely aligned with human visual processing, as evaluated against ten selected psycho/physiological visual properties.
Introduction
The introduction outlines the significance of a proposed evaluation program rooted in low-level visual psychophysics, which aims to complement higher-level proposals discussed in Bowers et al. (2023) and Biscione et al. (2024). Despite the novelty of this approach within the AI and machine learning communities, recent independent evaluations by researchers such as Cai et al. (2025) and Hammou et al. (2025) suggest a growing interest in similar methodologies. The authors assert that their evaluation program extends beyond the traditional parameters of luminance, color, and contrast masking, emphasizing a qualitative understanding of response curves alongside quantitative comparisons.
Furthermore, the introduction highlights the work of Rafal Mantiuk’s lab, which aligns with the focus on low-level psychophysics but differs in its emphasis on quantitative analysis. The proposed evaluation framework is designed to facilitate the identification of conceptual errors in deep learning models, enabling researchers to make informed modifications to model architectures based on the selected visual stimuli and their psychophysical properties. This approach aims to enhance the interpretability and effectiveness of AI systems by fostering a deeper understanding of their visual processing capabilities.
Methods
In this section, the authors address the complexities involved in evaluating artificial models of vision, particularly the choice between artificial physiology and psychophysics for experimental settings. They highlight the importance of selecting appropriate methods for visualizing model behavior, referencing established techniques from visual neuroscience, such as receptive field analysis and reverse correlation. The authors propose a methodology that involves generating stimuli—well-calibrated digital images that reflect key properties of low-level human vision—and using these to assess model responses. The evaluation can be performed through qualitative comparisons of response curves or quantitative measures like Pearson and rank correlations, depending on the availability of ground truth data.
The proposed methodology emphasizes the significance of selecting the correct layer from which to read model outputs, as different layers may correspond to distinct processing stages in the human visual system. The authors illustrate their approach using three illustrative networks: the BioMultiLayer model, PerceptNet, and Bio U-Net, each designed to capture various aspects of visual processing. They detail the experimental setup, including the use of 128 × 128 images and specific tests to measure model responses, thereby laying the groundwork for subsequent analyses of these models’ performance against established visual phenomena.
Results
In the “Results” section, the authors present both qualitative and quantitative analyses of model alignments with human visual behavior. The qualitative assessment, summarized in Table 2, indicates that the parametric model, BioMultiLayer, exhibits superior alignment with human responses, particularly in linear regions, due to its incorporation of sensitivities and center-surround mechanisms. Notably, the simulation of band-pass behavior in the context of the Blakemore and Campbell experiment highlights the model’s effectiveness in reproducing human-like responses. The alignment is evaluated through Pearson correlation and Kendall rank correlation, revealing that while the BioMultiLayer outperforms the nonparametric PerceptNet and Bio-U-Net, there are still areas, particularly in feature interaction, that require further refinement.
The quantitative analysis, illustrated in Figure 13, employs root mean square error (RMSE) alignment and correlation metrics to provide a detailed comparison of model performance across various properties. The BioMultiLayer achieves an average alignment score of 0.68 ± 0.14, significantly higher than the other models, which scored 0.26 ± 0.26 for PerceptNet and 0.20 ± 0.34 for Bio-U-Net. Statistical tests confirm the BioMultiLayer’s superior alignment with human behavior, while the other two models show no significant differences between them. The authors caution that while quantitative metrics offer valuable insights, they may overlook qualitative nuances, emphasizing the importance of visual assessments alongside numerical evaluations. The proposed methodology not only aids in understanding model-human alignment but also suggests pathways for enhancing artificial vision systems by mimicking human-like visual processing characteristics.
Discussion
In the discussion section of the paper, the authors critically evaluate the claims made by various researchers regarding the effectiveness of deep learning models in mimicking human visual processing. They emphasize the importance of skepticism in scientific inquiry, referencing historical figures like Tomaso Poggio and Horace Barlow, who have questioned the foundational frameworks of vision science. The authors propose that understanding the visual brain through AI techniques is more complex than previously assumed, advocating for more rigorous testing methods. They introduce a Turing-like test based on ten properties of low-level human vision, termed the “Decalogue,” to assess the performance of artificial models against human visual behavior.
The authors outline the structure of their paper, indicating that they will present a straightforward evaluation method using visual stimuli that reflect classical vision science principles. They plan to illustrate this method through the assessment of three different models, highlighting the necessity of ensuring that these models exhibit human-like adaptive properties. The discussion also raises broader issues regarding the computational goals of visual systems and the challenges of interpreting deep learning models as biologically plausible. The authors conclude by emphasizing the need for improved evaluation techniques to better understand how closely artificial models align with human visual processing, ultimately contributing to advancements in both AI and vision science.
