من الهيكل إلى التآزر: مسح لتطور نموذج إدراك الرؤية واللغة في نماذج اللغة الكبيرة متعددة الوسائط
From structure to synergy: A survey of vision-language perception paradigm evolution in multimodal large language models

شارك:
المجلة: Information Fusion، المجلد: 133
DOI: https://doi.org/10.1016/j.inffus.2026.104285
تاريخ النشر: 2026-03-07
المؤلف: Haoxiang Sun وآخرون
الموضوع الرئيسي: تطبيقات تعلم الآلة متعددة الوسائط

نظرة عامة

يقدم هذا القسم من ورقة البحث استعراضًا منهجيًا للإدراك الموحد للرؤية واللغة في نماذج اللغة الكبيرة متعددة الوسائط (MLLMs)، معالجًا فجوة ملحوظة في الأدبيات حيث غالبًا ما تعالج المراجعات الحالية الرؤية واللغة ككيانات منفصلة. يعرف المؤلفون إدراك MLLM كقدرة متكاملة تشبه الإدراك البشري ويقدمون تصنيفًا من خمس مراحل يتتبع تطور نماذج الإدراك في MLLMs. يشمل هذا التصنيف مراحل مختلفة، بما في ذلك تحسينات تركز على المشفر وأخرى تركز على المفكك، مما يؤدي إلى استراتيجيات معالجة أكثر ديناميكية وتكيفًا.

يسلط الاستعراض الضوء على تقارب ضبط التعليم، والتكيف، والتعلم المعزز كعوامل محورية في تطوير الجيل التالي من الوكلاء المتمركزين حول الإدراك. على الرغم من التقدم، يحدد المؤلفون تحديات كبيرة لا تزال قائمة، مثل الاعتماد على بيانات متعددة المهام المنسقة، ونقص المكافآت الدقيقة العامة، والتكاليف الحسابية العالية دون جدولة فعالة. يهدف المؤلفون إلى أن تكون نتائجهم بمثابة فهم أساسي وخارطة طريق قابلة للتنفيذ لدفع المزيد من الابتكار نحو تحقيق الذكاء العام الاصطناعي (AGI).

مقدمة

تؤكد مقدمة ورقة البحث على الدور الحاسم للمعرفة الدلالية في تحويل المعلومات الحسية إلى رؤى ذات مغزى. يقدم المؤلفون ثلاث مساهمات رئيسية في مجال إدراك الرؤية واللغة في نماذج اللغة الكبيرة متعددة الوسائط (MLLMs).

أولاً، يدعون أن هذا العمل هو أول استعراض منهجي يحلل إدراك الرؤية واللغة كقدرة متماسكة ضمن MLLMs، مقدمًا خارطة طريق منظمة للجهود البحثية المستقبلية. ثانيًا، يقدمون تصنيفًا من خمس مراحل يصنف تطور النماذج في إدراك MLLM، مصحوبًا بمراجعة شاملة للأعمال المهمة في كل مرحلة، مما يبرز ابتكاراتهم الرئيسية. أخيرًا، يتناول المؤلفون التحديات الناشئة في هذا المجال ويقترحون اتجاهات مستقبلية لتعزيز الذكاء متعدد الوسائط المتمركز حول الإدراك، بهدف توجيه جهود البحث اللاحقة.

الطرق

يستعرض قسم الطرق مختلف التقدمات في تقنيات التعلم المعزز (RL) المطبقة على مهام الإدراك البصري، مع التركيز بشكل خاص على اكتشاف العلاقات N-ary والتجزئة. يمتد البحث إلى طرق تحسين سياسة التعزيز العامة (GRPO) القابلة للتحقق ضمن إطار موحد لفهم العلاقات. تشمل المساهمات البارزة PixelThink، الذي يستخدم مخطط تفكير مدرك للكفاءة لتكييف طول التفكير بناءً على صعوبة المهمة، وVisionReasoner، الذي يدمج وظائف مكافأة مصممة خصيصًا للتفكير البصري متعدد المهام. بالإضافة إلى ذلك، يعزز ALToLLM تجزئة الكائنات من خلال مُرمز قناع بطول متكيف، بينما تستخدم Anomaly-R1 وOmniAD GRPO لتحسين اكتشاف الشذوذ الصناعي.

تُبرز الابتكارات الإضافية في منهجيات RL، مثل تحسين POPEN القائم على التفضيلات لتجزئة التفكير وPerPO الذي يتماشى مع الإدراك البشري للمهام التمييزية. يعزز VideoChat-TPO الإدراك الزمني المكاني في مقاطع الفيديو من خلال رؤوس محددة للمهام، ويطبق Vis-ReP التدريب الذاتي المعزز لتحسين تركيب البرامج البصرية. يناقش القسم أيضًا نظام Insight-V متعدد الوكلاء لتفكيك مهام التفكير البصري وإعادة صياغة SegAgent للتجزئة كعملية قرار متعددة الخطوات، مما يظهر نهجًا شاملاً لتحسين أداء النموذج عبر مهام الرؤية واللغة المختلفة.

المناقشة

في هذه المناقشة، يبرز المؤلفون تطور نماذج اللغة الكبيرة متعددة الوسائط (MLLMs) من الأطر المتمركزة حول النص إلى أنظمة الرؤية واللغة المتكاملة، مع التأكيد على قدراتها المحسنة في الإدراك. حققت MLLMs تقدمًا كبيرًا في المهام عبر الوسائط مثل وصف الصور، والتأسيس البصري، والإجابة على الأسئلة متعددة الوسائط، مدفوعةً بقدرتها على استخراج والتفكير في المعلومات البصرية بطريقة تشبه الإدراك البشري. ومع ذلك، غالبًا ما تقدم الأدبيات الحالية رؤية مجزأة، تركز إما على مهام الرؤية أو اللغة بشكل منفصل، مما يغفل جانب الإدراك الموحد الذي يعد حاسمًا لفهم الإمكانات الكاملة لـ MLLMs. يهدف هذا الاستعراض إلى سد تلك الفجوة من خلال تقديم تحليل منظم لتطور إدراك الرؤية واللغة في MLLMs، مصنفًا إياه إلى خمس مراحل متميزة تعكس الانتقال من تحسينات وحدات إلى أطر ديناميكية وموحدة.

يحدد المؤلفون نطاق استعراضهم، مركزين على الطرق التي تتضمن استفسارات باللغة الطبيعية ومدخلات بصرية، مع استبعاد الأعمال التي تعطي الأولوية للتفكير المنطقي أو تلك التي تعزز الأداء فقط من خلال توسيع الهيكل. يقدمون تصنيفًا من خمس مراحل يلتقط مسار قدرات الإدراك، بدءًا من تحسينات تركز على المشفر إلى استراتيجيات أكثر تعقيدًا تركز على المفكك والديناميكية. لا يوضح هذا الإطار فقط تقدم تحسينات الإدراك ولكنه يعمل أيضًا كخريطة طريق للبحث المستقبلي، مشجعًا على نهج أكثر تكاملاً لفهم التفاعل بين الرؤية واللغة في MLLMs. من خلال التركيز على الإدراك على مستوى المنطقة والحالة تحت توجيه اللغة الطبيعية، يهدف الاستعراض إلى تقديم نظرة شاملة حول كيفية تطور هياكل الإدراك، مما يؤدي في النهاية إلى نموذج إدراك موحد للرؤية واللغة.

القيود

يسلط قسم القيود الضوء على عدة تحديات حاسمة في تقييم نماذج اللغة الكبيرة متعددة الوسائط (MLLMs). أصبحت مجموعات البيانات التقليدية غير كافية بشكل متزايد لاحتياجات البحث الحالية، حيث تعاني المهام البديلة مثل الكشف الثابت والتجزئة من قيود التسمية البشرية، مما يعقد تقييم أداء MLLM. بالإضافة إلى ذلك، غالبًا ما تفتقر هذه المجموعات إلى الصلة بالتطبيقات العملية، وأداء النماذج الكبيرة على المعايير العامة يقترب من التشبع. يثير هذا أسئلة أساسية حول صلاحية المهام البديلة التقليدية وتمثيلها للفائدة في العالم الحقيقي.

يتناول القسم أيضًا معضلة “التعميم مقابل الضبط الدقيق”، مشيرًا إلى أن التطور السريع للنماذج متعددة الوسائط الأساسية قد قلل من فعالية الضبط الدقيق على مجموعات البيانات متوسطة الحجم، مثل COCO. غالبًا ما يجد الباحثون أن الانتظار للنسخة التالية من النموذج يؤدي إلى تحسينات أكبر في الأداء مقارنةً بالتحسينات المحددة للمهام. علاوة على ذلك، لا تزال تعريفات “متعددة الوسائط” مثيرة للجدل، حيث تؤدي التصنيفات الحالية حسب الوسائط إلى فهم مجزأ. بينما تم تحقيق نتائج مثيرة للإعجاب من خلال التدريب المسبق على اللغة والتوافق متعدد الوسائط، من المحتمل أن تواجه هذه الطريقة التوسعية قيودًا، مما يستلزم مزيدًا من التحقيق في وجود مساحة تمثيل موحدة عبر الوسائط. يختتم القسم بالإشارة إلى أن MLLMs تواجه اختناقات تقنية محددة سيتم توضيحها في الأقسام الفرعية التالية.

Journal: Information Fusion, Volume: 133
DOI: https://doi.org/10.1016/j.inffus.2026.104285
Publication Date: 2026-03-07
Author(s): Haoxiang Sun et al.
Primary Topic: Multimodal Machine Learning Applications

Overview

This research paper section presents a systematic survey of unified vision-language perception in Multimodal Large Language Models (MLLMs), addressing a notable gap in the literature where existing reviews often treat vision and language as separate entities. The authors define MLLM perception as an integrated capability akin to human perception and introduce a five-stage taxonomy that traces the evolution of perception paradigms in MLLMs. This taxonomy encompasses various phases, including encoder-centric and decoder-centric optimizations, leading to more dynamic and adaptive processing strategies.

The survey highlights the convergence of instruction tuning, adaptivity, and reinforcement learning as pivotal in developing the next generation of perception-centric agents. Despite the advancements, the authors identify significant challenges that persist, such as the reliance on curated multi-task data, the lack of general fine-grained rewards, and high computational costs without effective scheduling. The authors aim for their findings to serve as a foundational understanding and actionable roadmap to propel further innovation towards achieving artificial general intelligence (AGI).

Introduction

The introduction of the research paper emphasizes the critical role of semantic knowledge in transforming sensory information into meaningful insights. The authors present three main contributions to the field of vision-language perception in multimodal large language models (MLLMs).

Firstly, they claim this work is the first systematic survey that analyzes vision-language perception as a cohesive capability within MLLMs, offering a structured roadmap for future research endeavors. Secondly, they introduce a five-stage taxonomy that categorizes the evolution of paradigms in MLLM perception, accompanied by a comprehensive review of significant works at each stage, underscoring their key innovations. Lastly, the authors address emerging challenges in the field and propose future directions for enhancing perception-centric multimodal intelligence, aiming to guide subsequent research efforts.

Methods

The methods section outlines various advancements in reinforcement learning (RL) techniques applied to visual perception tasks, particularly focusing on N-ary relation detection and segmentation. The research extends verifiable-reward Generalized Reinforcement Policy Optimization (GRPO) methods within a unified framework for relation comprehension. Notable contributions include PixelThink, which employs an efficiency-aware reasoning scheme to adapt reasoning length based on task difficulty, and VisionReasoner, which integrates tailored reward functions for multi-task visual reasoning. Additionally, ALToLLM enhances object segmentation through an adaptive-length mask tokenizer, while Anomaly-R1 and OmniAD utilize GRPO to improve industrial anomaly detection.

Further innovations in RL methodologies are highlighted, such as POPEN’s preference-based optimization for reasoning segmentation and PerPO’s alignment with human perception for discriminative tasks. VideoChat-TPO enhances spatio-temporal perception in videos through task-specific heads, and Vis-ReP applies reinforced self-training to improve visual program synthesis. The section also discusses Insight-V’s multi-agent system for decomposing visual reasoning tasks and SegAgent’s reformulation of segmentation as a multistep decision process, demonstrating a comprehensive approach to optimizing model performance across various vision-language tasks.

Discussion

In this discussion, the authors highlight the evolution of Multimodal Large Language Models (MLLMs) from text-centric frameworks to integrated vision-language systems, emphasizing their enhanced perception capabilities. MLLMs have made significant strides in cross-modal tasks such as image captioning, visual grounding, and multimodal question answering, driven by their ability to extract and reason about visual information in a manner akin to human cognition. However, existing literature often presents a fragmented view, focusing either on vision or language tasks separately, which neglects the unified perception aspect critical for understanding MLLMs’ full potential. This survey aims to fill that gap by providing a structured analysis of the evolution of vision-language perception in MLLMs, categorizing it into five distinct stages that reflect the transition from modular optimization to dynamic and unified frameworks.

The authors delineate the scope of their survey, concentrating on methods that involve natural-language queries and visual inputs, while explicitly excluding works that prioritize logical reasoning or those that solely enhance performance through architectural scaling. They introduce a five-stage taxonomy that captures the trajectory of perception capabilities, starting from encoder-centric optimizations to more sophisticated decoder-centric and dynamic strategies. This framework not only illustrates the progression of perception enhancements but also serves as a roadmap for future research, encouraging a more integrated approach to understanding the interplay between vision and language in MLLMs. By focusing on region- and instance-level perception under natural-language guidance, the survey aims to provide a comprehensive overview of how perception architectures have evolved, ultimately leading to a unified vision-language perception paradigm.

Limitations

The section on limitations highlights several critical challenges in the evaluation of Multimodal Large Language Models (MLLMs). Traditional datasets are increasingly inadequate for current research needs, as proxy tasks like static detection and segmentation suffer from the limitations of human labeling, complicating the assessment of MLLM performance. Additionally, these datasets often lack relevance to practical applications, and the performance of large models on general-purpose benchmarks is nearing saturation. This raises essential questions about the validity of traditional proxy tasks and their representation of real-world utility.

The section also addresses the “Generalization vs. Fine-tuning” dilemma, noting that the rapid evolution of foundational multimodal models has diminished the effectiveness of fine-tuning on medium-scale datasets, such as COCO. Researchers often find that waiting for the next model iteration yields greater performance improvements than task-specific optimizations. Furthermore, the definition of “Multimodal” remains contentious, as current categorizations by modality lead to a fragmented understanding. While impressive results have been achieved through language pre-training and multimodal alignment, this scaling approach is likely to encounter limitations, necessitating further investigation into the existence of a unified representation space across modalities. The section concludes by indicating that MLLMs face specific technical bottlenecks that will be elaborated upon in subsequent subsections.

شارك: