DOI: https://doi.org/10.1038/s41598-025-33794-x
PMID: https://pubmed.ncbi.nlm.nih.gov/41530247
تاريخ النشر: 2026-01-13
المؤلف: Basab Nath وآخرون
الموضوع الرئيسي: تقنيات معالجة اللغة الطبيعية
نظرة عامة
تقدم هذه البحث إطار عمل جديد لترجمة الآلات العصبية (NMT) تستهدف اللغات الهندية ذات الموارد المنخفضة، وبالتحديد الأسامية والبودو. قام المؤلفون بضبط نموذج المحول متعدد اللغات ثنائي الاتجاه والانحداري التلقائي (mBART50) على مجموعات بيانات الأسامية-الإنجليزية والبودو-الإنجليزية، على الرغم من عدم تضمين هذه اللغات في تدريب النموذج المسبق. قدموا نهجًا من مرحلتين يدمج هياكل مزيج الخبراء النادرة مع تقطير المعرفة عبر اللغات، مما أدى إلى نموذج طالب مدمج يحتوي على 400 مليون معلمة. يحقق هذا النموذج جودة ترجمة قابلة للمقارنة مع معلمه الذي يحتوي على 1.3 مليار معلمة، مع فرق قدره 0.7 نقطة BLEU في الأسامية-الإنجليزية (34.5 مقابل 35.2) وفرق قدره 0.8 نقطة في البودو-الإنجليزية (31.2 مقابل 32.0)، بينما يعزز بشكل كبير سرعة الاستدلال بحوالي 280% على الأجهزة ذات الجودة الاستهلاكية.
تشير النتائج إلى أن الطريقة المقترحة تحتفظ بشكل فعال بأكثر من 97% من أداء BLEU للمعلم بينما تتيح نشرًا أسرع، مما يجعلها مناسبة للتطبيقات الواقعية مثل الأجهزة المحمولة والمنصات التعليمية. أثبتت مجموعة من تقطير المعرفة على مستوى التسلسل وهندسة مزيج الخبراء النادرة الهجينة أنها مفيدة، كما أكدت الدراسات التجريبية والتحليلات النوعية التي أبرزت دقة المعاني في الترجمات. يقترح المؤلفون أن إطار عملهم لا يسهل فقط الترجمة الفعالة للغات غير الممثلة، بل يضع أيضًا الأساس للتوسعات المستقبلية إلى لغات أخرى ذات موارد منخفضة ضمن عائلتي الهندو-أرية والتبتو-بورمانية، بالإضافة إلى التعديلات المحتملة للمهام الخاصة بالمجالات.
مقدمة
تناقش مقدمة ورقة البحث التقدمات في ترجمة الآلات العصبية (NMT)، وخاصة للغات ذات الموارد العالية، التي تسهلها نماذج متعددة اللغات مدربة مسبقًا مثل mBART. على الرغم من هذه النجاحات، تظل العديد من اللغات الهندية، مثل الأسامية والبودو، ذات موارد منخفضة، تواجه تحديات بسبب نقص مجموعات البيانات المتوازية والتمثيل غير الكافي في أنظمة الترجمة الحالية. لقد حققت المبادرات الأخيرة، مثل IndicTrans، خطوات كبيرة في معالجة هذه القضايا من خلال استخدام محول متعدد اللغات يحتوي على 1.3 مليار معلمة ويشمل جميع اللغات الهندية المجدولة البالغ عددها 22، مما أدى إلى تحسينات ملحوظة في درجات BLEU. ومع ذلك، فإن أحجام النماذج الكبيرة تعيق النشر على الأجهزة.
لتخفيف المطالب الحاسوبية والذاكرة لهذه النماذج، يتم استكشاف تقنيات مثل تقطير المعرفة وهياكل مزيج الخبراء النادرة (MoE). يسمح إطار MoE بالحساب الشرطي من خلال تدريب عدة شبكات متخصصة، أو “خبراء”، بالتوازي، مع شبكة بوابة تختار مجموعة فرعية من هؤلاء الخبراء لكل إدخال. تتيح هذه الطريقة توسيع النماذج إلى مليارات المعلمات مع تفعيل جزء فقط أثناء الاستدلال، وبالتالي تحقيق توازن بين السعة والكفاءة الحاسوبية. تسلط الورقة الضوء على ضرورة معالجة التحديات الفريدة التي تطرحها اللغات الهندية ذات الموارد المنخفضة، خاصة في سياق تطبيقات الترجمة في الوقت الحقيقي وعلى الأجهزة.
طرق
في هذا القسم، يحدد المؤلفون إطارهم المنهجي، الذي يشمل اختيار لغات البرمجة ومجموعات البيانات المستخدمة في الدراسة. يوضحون بنية النموذج المقترح، مع التركيز على تصميمه ووظيفته. بالإضافة إلى ذلك، تتضمن المنهجية وصفًا لعملية تقطير المعرفة التي تم تنفيذها لتعزيز أداء النموذج، إلى جانب إعداد التدريب الذي تم استخدامه. أخيرًا، يحدد المؤلفون مقاييس التقييم المستخدمة لقياس كل من فعالية وكفاءة نماذجهم، مما يضمن تقييمًا شاملاً لنهجهم.
نتائج
يقيم قسم النتائج أداء نموذج ترجمة الآلات العصبية (NMT) المدمج، الذي تم تقطيره من نموذج معلم كبير باستخدام إطار عمل تقطير المعرفة عبر اللغات من مرحلتين. تم إجراء تقييمات كمية باستخدام مقاييس تلقائية قياسية، وخاصة درجات BLEU وChrF، لقياس جودة الترجمة. كما درست الدراسة كفاءة النموذج الطالب الحاسوبية، مع التركيز على سرعة الاستدلال وبصمة النموذج.
بالإضافة إلى المقاييس الكمية، تم تقديم مقارنات نوعية لتقييم دقة الترجمة. سلط التحليل الضوء على قدرة النموذج الطالب على التعميم بشكل فعال عبر لغات ذات موارد منخفضة مختلفة، بهدف تحديد ما إذا كان يمكنه الحفاظ على جودة الترجمة مع تحقيق تخفيضات كبيرة في زمن الاستدلال واستخدام الذاكرة. بشكل عام، تشير النتائج إلى أن النموذج المقترح يوازن بنجاح بين أداء الترجمة وكفاءة الحوسبة.
نقاش
يسلط النقاش الضوء على التقدمات التي تم إحرازها في الترجمة الثنائية للغات الهندية، مع التركيز بشكل خاص على الأسامية والبودو، من خلال ضبط نموذج mBART50. على الرغم من نجاح المبادرات مثل Samanantar وIndicTrans، فإن الحجم الكبير للمعلمات في النماذج الحالية يحد من نشرها في البيئات ذات الموارد المحدودة. يقترح المؤلفون نهجًا من مرحلتين يقوم أولاً بضبط نموذج mBART50 المدرب مسبقًا على بيانات الأسامية-الإنجليزية والبودو-الإنجليزية، مما يمكّن من نقل فعال عبر اللغات من لغات ذات صلة. يعد هذا الضبط ضروريًا لتكييف النموذج لإنتاج تمثيلات ذات مغزى لهذه الأزواج اللغوية ذات الموارد المنخفضة، والتي تظهر اختلافات لغوية كبيرة.
تتضمن المرحلة الثانية تقطير النموذج المضبوط إلى نموذج مزيج خبراء (MoE) مدمج يحتوي على 400 مليون معلمة، مما يحقق توازنًا بين جودة الترجمة وكفاءة الحوسبة. يحتفظ النموذج الطالب بأكثر من 96% من درجة BLEU للمعلم بينما يحقق تسريعًا بنسبة 280% في زمن الاستدلال، مما يوضح ملاءمته للتطبيقات في الوقت الحقيقي على الأجهزة الطرفية. تمثل هذه المجموعة المبتكرة من تقطير المعرفة وهندسة مزيج الخبراء النادرة خطوة كبيرة نحو تعزيز قدرات الترجمة للغات غير الممثلة، مع معالجة كل من تحديات الجودة والنشر في البيئات ذات الموارد المنخفضة.
DOI: https://doi.org/10.1038/s41598-025-33794-x
PMID: https://pubmed.ncbi.nlm.nih.gov/41530247
Publication Date: 2026-01-13
Author(s): Basab Nath et al.
Primary Topic: Natural Language Processing Techniques
Overview
This research presents a novel framework for neural machine translation (NMT) targeting low-resource Indic languages, specifically Assamese and Bodo. The authors fine-tuned the Multilingual Bidirectional and Auto-Regressive Transformer (mBART50) model on Assamese-English and Bodo-English corpora, despite these languages not being included in the model’s pretraining. They introduced a two-stage approach that integrates sparse Mixture-of-Experts (MoE) architectures with cross-lingual knowledge distillation, resulting in a compact 400-million-parameter student model. This model achieves translation quality comparable to its 1.3-billion-parameter teacher, with only a 0.7 BLEU point difference in Assamese-English (34.5 vs. 35.2) and a 0.8 point difference in Bodo-English (31.2 vs. 32.0), while significantly enhancing inference speed by approximately 280% on consumer-grade hardware.
The findings indicate that the proposed method effectively retains over 97% of the teacher’s BLEU performance while enabling faster deployment, making it suitable for real-world applications such as mobile devices and educational platforms. The combination of sequence-level knowledge distillation and a hybrid sparse MoE architecture proved beneficial, as confirmed by ablation studies and qualitative analyses that highlighted the semantic fidelity of the translations. The authors suggest that their framework not only facilitates efficient translation for underrepresented languages but also lays the groundwork for future expansions to other low-resource languages within the Indo-Aryan and Tibeto-Burman families, as well as potential adaptations for domain-specific tasks.
Introduction
The introduction of the research paper discusses the advancements in neural machine translation (NMT), particularly for high-resource languages, facilitated by large multilingual pre-trained models like mBART. Despite these successes, many Indic languages, such as Assamese and Bodo, remain under-resourced, facing challenges due to limited parallel corpora and insufficient representation in existing translation systems. Recent initiatives, such as IndicTrans, have made strides in addressing these issues by employing a 1.3 billion parameter multilingual transformer that encompasses all 22 scheduled Indian languages, resulting in notable improvements in BLEU scores. However, the large model sizes hinder on-device deployment.
To mitigate the computational and memory demands of these models, techniques such as knowledge distillation and sparse Mixture-of-Experts (MoE) architectures are explored. The MoE framework allows for conditional computation by training multiple specialized networks, or “experts,” in parallel, with a gating network selecting a subset of these experts for each input. This method enables the scaling of models to billions of parameters while activating only a fraction during inference, thus balancing capacity and computational efficiency. The paper highlights the necessity of addressing the unique challenges posed by low-resource Indic languages, particularly in the context of real-time and on-device translation applications.
Methods
In this section, the authors outline their methodological framework, which encompasses the selection of programming languages and datasets utilized in the study. They detail the architecture of the proposed model, emphasizing its design and functionality. Additionally, the methodology includes a description of the knowledge distillation process implemented to enhance model performance, alongside the training setup that was employed. Finally, the authors specify the evaluation metrics used to measure both the effectiveness and efficiency of their models, ensuring a comprehensive assessment of their approach.
Results
The results section evaluates the performance of a compact neural machine translation (NMT) model, which was distilled from a large-scale teacher model using a two-stage cross-lingual knowledge distillation framework. Quantitative assessments were conducted using standard automatic metrics, specifically BLEU and ChrF scores, to measure translation quality. The study also examined the computational efficiency of the student model, focusing on inference speed and model footprint.
In addition to quantitative metrics, qualitative comparisons were provided to assess translation fidelity. The analysis highlighted the student model’s capability to generalize effectively across various low-resource languages, aiming to determine whether it could maintain translation quality while achieving significant reductions in inference latency and memory usage. Overall, the findings suggest that the proposed model successfully balances translation performance with computational efficiency.
Discussion
The discussion highlights the advancements made in bilingual translation for Indian languages, particularly focusing on Assamese and Bodo, through the fine-tuning of the mBART50 model. Despite the success of initiatives like Samanantar and IndicTrans, the large parameter size of existing models limits their deployment in resource-constrained environments. The authors propose a two-stage approach that first fine-tunes a pre-trained mBART50 model on Assamese-English and Bodo-English data, enabling effective cross-lingual transfer from related languages. This fine-tuning is crucial for adapting the model to produce meaningful representations for these low-resource language pairs, which exhibit significant linguistic differences.
The second stage involves distilling the fine-tuned model into a more compact 400M-parameter Mixture-of-Experts (MoE) model, achieving a balance between translation quality and computational efficiency. The student model retains over 96% of the teacher’s BLEU score while achieving a 280% speedup in inference time, demonstrating its suitability for real-time applications on edge devices. This innovative combination of knowledge distillation and sparse MoE architecture represents a significant step toward enhancing translation capabilities for underrepresented languages, addressing both quality and deployment challenges in low-resource settings.
