الحرباء: مسرع شبكة الالتفاف الزمني الخالي من المضاعفات للتعلم القليل والداوم من البيانات التسلسلية
Chameleon: A Multiplier-Free Temporal Convolutional Network Accelerator for End-to-End Few-Shot and Continual Learning from Sequential Data

شارك:
المجلة: IEEE Journal of Solid-State Circuits، المجلد: 61، العدد: 7
DOI: https://doi.org/10.1109/jssc.2025.3645640
تاريخ النشر: 2026-01-14
المؤلف: Douwe den Blanken وآخرون
الموضوع الرئيسي: تكييف المجال والتعلم من عدد قليل من الأمثلة

نظرة عامة

تقدم البحث “Chameleon”، وهو هيكل مبتكر مصمم للتعلم على الجهاز عند الحافة، حيث يعالج تحديات انخفاض الكمون، والتخصيص الخاص، والتعلم الفعال من بيانات تسلسلية محدودة. يدمج “Chameleon” إطار عمل موحد للتعلم والاستدلال يدعم التعلم من خلال عدد قليل من الأمثلة (FSL) والتعلم المستمر (CL) مع الحد الأدنى من تكلفة المساحة (0.5%) على منطق الاستدلال الحالي. يستخدم هذا الهيكل الشبكات التلافيفية الزمنية (TCNs) لالتقاط الاعتمادات الزمنية الطويلة بفعالية، مما يمثل أول تنفيذ ناجح للتعلم من خلال عدد قليل من الأمثلة والتعلم المستمر على الرقاقة للبيانات التسلسلية، تم إثباته بشكل خاص من خلال الاستدلال على صوت خام بتردد 16 كيلو هرتز.

يتميز “Chameleon” بمصفوفة حسابية ثنائية الوضع، خالية من المضاعفات، يمكن أن تتطابق إما مع استهلاك الطاقة لمسرعات اكتشاف الكلمات الرئيسية (KWS) التي تركز فقط على الاستدلال أو تحقق أداءً ذروياً أعلى بمقدار 4.3× من GOPS. تم تصنيعه بتقنية CMOS 40 نانومتر، يحقق “Chameleon” معايير دقة ملحوظة، بما في ذلك 96.8% لمهام 5-way 1-shot و98.8% لمهام 5-way 5-shot على مجموعة بيانات Omniglot، ودقة نهائية تبلغ 82.2% للتعلم المستمر عبر 250 فئة مع 10 لقطات. بالإضافة إلى ذلك، يحافظ على دقة استدلال عالية تبلغ 93.3% على مجموعة بيانات أوامر الكلام من Google المكونة من 12 فئة أثناء العمل ضمن ميزانية طاقة صارمة تبلغ 3.1 ميكرو واط.

مقدمة

تناقش مقدمة الورقة الحاجة المتزايدة لنشر نماذج التعلم العميق على الأجهزة الطرفية ذات الموارد المحدودة، لا سيما ضمن نظام إنترنت الأشياء (IoT). تركز أجهزة الذكاء الاصطناعي الحالية على الحافة بشكل أساسي على الاستدلال باستخدام نماذج مدربة مسبقًا، مما يحد من قدرتها على التكيف مع البيانات والميزات الجديدة، وبالتالي يؤثر على موثوقيتها مع مرور الوقت. بينما يمكن أن يؤدي إعادة التدريب المستضاف على السحابة إلى تحديث النماذج ببيانات جديدة، فإنه يعوقه مشاكل الاتصال، وقيود الكمون، ومخاوف الخصوصية. على العكس من ذلك، فإن التدريب على الجهاز من الصفر يتطلب طاقة كبيرة ويحتاج إلى مجموعات بيانات موسومة كبيرة، والتي غالبًا ما تكون غير عملية في سيناريوهات الحافة.

لمعالجة هذه التحديات، تقدم الورقة “Chameleon”، وهي بنية جديدة تمكّن التعلم الفعال من خلال عدد قليل من الأمثلة (FSL) والتعلم المستمر (CL) للبيانات التسلسلية دون المساس بأداء الاستدلال. تشمل الابتكارات الرئيسية دمج التعلم ضمن عملية الاستدلال باستخدام الشبكات النموذجية، مما يقلل من تكلفة المساحة والكمون بينما يحقق دقة عالية في مهام FSL وCL. بالإضافة إلى ذلك، يعزز استخدام الشبكات التلافيفية الزمنية قدرة النموذج على التقاط الاعتمادات بعيدة المدى في البيانات التسلسلية، وتعمل مصفوفة عناصر المعالجة ثنائية الوضع على تحسين استهلاك الطاقة لتطبيقات الحافة القصوى. تؤكد الورقة على أهمية تحقيق التوازن بين الكفاءة والمرونة في حلول الذكاء الاصطناعي على الحافة وتوفر موارد مفتوحة المصدر لتسهيل المزيد من البحث والتطوير.

مناقشة

تحدد قسم المناقشة في الورقة إطار عمل للتعلم الميتا مصمم للتعلم من خلال عدد قليل من الأمثلة (FSL) وآثاره على النشر على الرقاقة. في التعلم الخاضع للإشراف التقليدي، تشترك مجموعات بيانات التدريب والاختبار في الفئات ولكن ليس في الأمثلة، بينما يخصص التعلم الميتا الفئات الكاملة إما للتدريب أو للاختبار، مما يسمح بمجموعات فئات متميزة في كل منهما. يستخدم الإطار مجموعة تدريب ميتا من المهام لتطوير شبكة عصبية عميقة (DNN) قادرة على التكيف مع المهام غير المرئية مع الحد الأدنى من الأمثلة. تعتبر مجموعات الدعم والاستعلام حاسمة للتدريب وتقييم الأداء، على التوالي، مع هيكل مشابه لمرحلة الاختبار الميتا.

تصنف الورقة طرق التعلم الميتا الحالية إلى ثلاث مجموعات: طرق تهيئة المعلمات، والنماذج الأمامية، وطرق تعلم القياس، كل منها له مزايا وتحديات متميزة للنشر على الرقاقة. تتيح طرق تهيئة المعلمات، مثل التعلم الميتا المستقل عن النموذج (MAML)، التكيف السريع ولكنها تتطلب نزول التدرج على الرقاقة، مما قد يكون مكلفًا حسابيًا. تلغي النماذج الأمامية الحاجة إلى تحديثات التدرج ولكن قد تقدم تكلفة معلمات كبيرة. تولد طرق تعلم القياس، مثل الشبكات النموذجية، تمثيلات متجهية للتصنيف دون تحديثات التدرج، مما يجعلها فعالة للاستخدام على الرقاقة. يقترح المؤلفون بنية نظام “Chameleon” على الرقاقة، والتي تدمج هذه الاستراتيجيات التعليمية في إطار عمل موحد، مما يحسن من التعلم والاستدلال مع تقليل تكلفة الموارد. هذه البنية مناسبة بشكل خاص لمعالجة البيانات التسلسلية، مستفيدة من الشبكات التلافيفية الزمنية (TCNs) لالتقاط الاعتمادات بعيدة المدى بكفاءة، وبالتالي معالجة تحديات نشر التعلم من خلال عدد قليل من الأمثلة والتعلم المستمر في التطبيقات الواقعية.

Journal: IEEE Journal of Solid-State Circuits, Volume: 61, Issue: 7
DOI: https://doi.org/10.1109/jssc.2025.3645640
Publication Date: 2026-01-14
Author(s): Douwe den Blanken et al.
Primary Topic: Domain Adaptation and Few-Shot Learning

Overview

The research presents Chameleon, an innovative architecture designed for on-device learning at the edge, addressing the challenges of low-latency, private personalization, and efficient learning from limited sequential data. Chameleon integrates a unified learning and inference framework that supports few-shot learning (FSL) and continual learning (CL) with minimal area overhead (0.5%) to the existing inference logic. This architecture employs temporal convolutional networks (TCNs) to effectively capture long temporal dependencies, marking the first successful implementation of end-to-end on-chip FSL and CL for sequential data, specifically demonstrated through inference on 16-kHz raw audio.

Chameleon features a dual-mode, multiplier-free compute array that can either match the power consumption of leading inference-only keyword spotting (KWS) accelerators or achieve a peak performance of 4.3× higher GOPS. Fabricated in 40-nm CMOS technology, Chameleon achieves remarkable accuracy benchmarks, including 96.8% for 5-way 1-shot and 98.8% for 5-way 5-shot tasks on the Omniglot dataset, and an 82.2% final accuracy for CL across 250 classes with 10 shots. Additionally, it maintains a high inference accuracy of 93.3% on the 12-class Google Speech Commands dataset while operating within a stringent power budget of 3.1 µW.

Introduction

The introduction of the paper discusses the growing need for deploying deep learning models on resource-constrained edge devices, particularly within the Internet-of-Things (IoT) ecosystem. Current edge AI devices primarily focus on inference using pre-trained models, which limits their adaptability to new data and features, thereby affecting their reliability over time. While cloud-hosted retraining can update models with new data, it is hindered by connectivity issues, latency constraints, and privacy concerns. Conversely, on-device training from scratch is energy-intensive and requires large labeled datasets, which are often impractical in edge scenarios.

To address these challenges, the paper introduces “Chameleon,” a novel architecture that enables efficient on-chip few-shot learning (FSL) and continual learning (CL) for sequential data without compromising inference performance. Key innovations include the integration of learning within the inference process using prototypical networks, which minimizes area overhead and latency while achieving high accuracy in FSL and CL tasks. Additionally, the use of temporal convolutional networks enhances the model’s ability to capture long-range dependencies in sequential data, and a dual-mode processing element array optimizes power consumption for extreme-edge applications. The paper emphasizes the importance of balancing efficiency and versatility in edge AI solutions and provides open-source resources to facilitate further research and development.

Discussion

The discussion section of the paper outlines a meta-learning framework tailored for few-shot learning (FSL) and its implications for on-chip deployment. In traditional supervised learning, training and testing datasets share classes but not examples, while meta-learning assigns entire classes to either training or testing splits, allowing for distinct class sets in each. The framework utilizes a meta-training set of tasks to develop a deep neural network (DNN) capable of adapting to unseen tasks with minimal examples. The support and query sets are crucial for training and performance evaluation, respectively, with the meta-testing phase similarly structured.

The paper categorizes existing meta-learning methods into three groups: parameter initialization, feed-forward, and metric-learning approaches, each with distinct advantages and challenges for on-chip deployment. Parameter initialization methods, like Model-Agnostic Meta-Learning (MAML), enable rapid adaptation but require on-chip gradient descent, which can be computationally intensive. Feed-forward models eliminate the need for gradient updates but may introduce significant parameter overhead. Metric-learning methods, such as prototypical networks, generate vector representations for classification without gradient updates, making them efficient for on-chip use. The authors propose the Chameleon system-on-chip architecture, which integrates these learning strategies into a unified framework, optimizing for both learning and inference while minimizing resource overhead. This architecture is particularly suited for sequential data processing, leveraging temporal convolutional networks (TCNs) to capture long-range dependencies efficiently, thereby addressing the challenges of deploying FSL and continual learning in real-world applications.

شارك: