محو الأمية الخوارزمية: تقنيات الذكاء الاصطناعي التوليدي لمكتبي البيانات
ALGORITHMIC LITERACY: Generative Artificial Intelligence Technologies for Data Librarians

المجلة: ICST Transactions on Scalable Information Systems، المجلد: 11، العدد: 2
DOI: https://doi.org/10.4108/eetsis.4067
تاريخ النشر: 2024-01-11
المؤلف: Alexandre Ribas Semeler وآخرون
الموضوع الرئيسي: الحوسبة العلمية وإدارة البيانات

نظرة عامة

يتناول قسم ورقة البحث دمج تقنيات الذكاء الاصطناعي (AI) في خدمات المكتبات الأكاديمية، مع التركيز بشكل خاص على دور OpenAI Codex في مساعدة أمناء البيانات. تهدف الدراسة إلى استكشاف كيفية تسهيل الذكاء الاصطناعي لتوليد نصوص برمجية بلغة بايثون لمهام استخراج البيانات، مع معالجة التحديات التي تطرحها البيانات الضخمة. من خلال الاستفادة من الذكاء الاصطناعي، يمكن لأمناء البيانات تقديم أوصاف باللغة الطبيعية لمشاكل البرمجة، والتي يقوم OpenAI Codex بعد ذلك بترجمتها إلى كود بايثون وظيفي، مما يعزز قدرتهم على إدارة واستخراج البيانات من مصادر متنوعة.

تشمل المنهجية التعاون بين أمناء البيانات ووكلاء الذكاء الاصطناعي، باستخدام أدوات مثل بيئة تطوير بايثون المتكاملة PyCharm وملحقات مثل Machinet AI وBito AI. تعرض النتائج خمسة سيناريوهات محددة لاستخراج البيانات من الويب، مما يوضح فعالية النصوص البرمجية التي تم إنشاؤها بواسطة الذكاء الاصطناعي في استخراج البيانات، وحساب المقاييس، وحفظ النتائج. تؤكد الخاتمة على إمكانية تمكين الذكاء الاصطناعي لأمناء البيانات غير المبرمجين، مما يمكنهم من التنقل في بيئات البيانات المعقدة وتحسين خدمات المكتبة من خلال توليد الكود تلقائيًا، مما يعزز في النهاية قدرتهم على التعامل مع مصادر البيانات المتنوعة في المشهد الرقمي.

مقدمة

تناقش مقدمة هذه الورقة البحثية دمج تقنيات الذكاء الاصطناعي (AI) في المكتبات الأكاديمية، مع التأكيد على العلاقة التبادلية بين أدوات الذكاء الاصطناعي والدور المتطور لأمناء البيانات. يمكّن الذكاء الاصطناعي، وخاصة من خلال منصات مثل نموذج OpenAI المدرب مسبقًا (GPT)، الأمناء من إدارة وتفسير كميات هائلة من البيانات، مما يعالج التحديات التي تطرحها الفيضانات البيانية. تسلط الورقة الضوء على كيفية تحسين أدوات الذكاء الاصطناعي لخدمات المكتبة من خلال أتمتة المهام، وتحسين إدارة البيانات، وتسهيل إنشاء خدمات جديدة، مثل روبوتات الدردشة وأنظمة جمع البيانات.

يقترح المؤلفون إطارًا تجريبيًا حيث يمكن لأمناء البيانات، حتى بدون خبرة برمجية، استخدام أدوات مدعومة بالذكاء الاصطناعي لتوليد كود بلغة بايثون بناءً على تعليمات باللغة الطبيعية. تهدف هذه المقاربة إلى تعزيز معرفة الأمناء بالخوارزميات ومهارات القرصنة، مما يمكنهم من التفاعل بفعالية مع تقنيات الذكاء الاصطناعي. تحدد الدراسة تطبيقات محددة، بما في ذلك استخراج البيانات من الويب للبيانات الببليومترية من مصادر مثل ORCID وGoogle Scholar، وتقدم دراسة حالة منظمة تستكشف الكفاءات المطلوبة لأمناء البيانات في عصر الذكاء الاصطناعي. في النهاية، تؤكد الأبحاث على إمكانية تحويل الذكاء الاصطناعي لخدمات المكتبة وتمكين الأمناء من الاستفادة من التكنولوجيا لحلول مبتكرة في إدارة البيانات.

النتائج

يستعرض قسم النتائج مبادرة عامة تهدف إلى إنشاء نصوص ومجموعات بيانات تم إنشاؤها بواسطة تقنيات الذكاء الاصطناعي خصيصًا لأمناء البيانات غير المبرمجين. تؤكد هذه المشروع على إمكانية الوصول وسهولة الاستخدام، مما يضمن أن الموارد المطورة يمكن استخدامها بفعالية من قبل الأفراد الذين ليس لديهم خبرة برمجية.

تكون جميع البيانات والنصوص والمجموعات التي تم جمعها متاحة للجمهور وفقًا لخطة إدارة بيانات البحث (PGD) وتكون مرخصة بموجب CC BY 4.0. يمكن الوصول إلى هذه الموارد من خلال مستودع Mendeley، مما يعزز العلوم المفتوحة ويسهل التفاعل الأوسع مع أدوات الذكاء الاصطناعي في إدارة البيانات.

المناقشة

تؤكد المناقشة حول إدارة البيانات على الدور المتطور لأمناء البيانات، مع تسليط الضوء على ضرورة المعرفة الرقمية والحسابية والبيانية في ممارستهم. مع تحول المنهجيات المعتمدة على البيانات إلى جزء لا يتجزأ من البحث الأكاديمي، يُحث الأمناء على التكيف من خلال اكتساب مهارات تتعلق بإدارة البيانات، والتنسيق، والتصور. تشير الأدبيات إلى أن معرفة البيانات مرتبطة ارتباطًا وثيقًا بمعرفة المعلومات، مما يتطلب فهمًا شاملاً لجودة البيانات والاستخدام الأخلاقي. علاوة على ذلك، أدت ظهور البيانات الضخمة إلى تطوير مفاهيم جديدة للمعرفة، بما في ذلك معرفة الخوارزميات، والتي تعتبر ضرورية للتنقل في تقنيات الذكاء الاصطناعي وفهم آثار اتخاذ القرار المدفوع بالخوارزميات.

تُعرف معرفة الخوارزميات بأنها القدرة على فهم وتطبيق الخوارزميات في سياقات مختلفة، وهو أمر أصبح أكثر أهمية لأمناء البيانات. تشمل هذه المعرفة ليس فقط مهارات البرمجة ولكن أيضًا القدرة على التفاعل بشكل نقدي مع الخوارزميات وتأثيراتها الاجتماعية. تؤكد الورقة على أهمية دمج معرفة الخوارزميات في تعليم المكتبات، مما يمكّن الأمناء من الاستفادة من أدوات الذكاء الاصطناعي بفعالية. يدعو المؤلفون إلى نهج تعاوني حيث يستخدم الأمناء أدوات مدعومة بالذكاء الاصطناعي، مثل OpenAI Codex، لتعزيز قدراتهم البرمجية لمهام مثل استخراج البيانات من الويب، مما يعزز قوة عاملة أكثر معرفة بالبيانات قادرة على مواجهة التحديات المعاصرة في إدارة المعلومات.

Journal: ICST Transactions on Scalable Information Systems, Volume: 11, Issue: 2
DOI: https://doi.org/10.4108/eetsis.4067
Publication Date: 2024-01-11
Author(s): Alexandre Ribas Semeler et al.
Primary Topic: Scientific Computing and Data Management

Overview

The research paper section discusses the integration of artificial intelligence (AI) technologies into academic library services, particularly focusing on the role of the OpenAI Codex in assisting data librarians. The study aims to explore how AI can facilitate the generation of Python code scripts for web scraping tasks, addressing the challenges posed by big data. By leveraging AI, data librarians can provide natural language descriptions of programming problems, which the OpenAI Codex then translates into functional Python code, enhancing their ability to manage and extract data from various sources.

The methodology involves collaboration between data librarians and AI agents, utilizing tools such as the PyCharm integrated development environment and plugins like Machinet AI and Bito AI. The results showcase five specific web-scraping scenarios, illustrating the effectiveness of AI-generated scripts in extracting data, calculating metrics, and saving results. The conclusion emphasizes the potential of AI to empower non-programming data librarians, enabling them to navigate complex data environments and improve library services through automated code generation, ultimately enhancing their capacity to handle diverse data sources in the digital landscape.

Introduction

The introduction of this research paper discusses the integration of artificial intelligence (AI) technologies into academic libraries, emphasizing the symbiotic relationship between AI tools and the evolving role of data librarians. AI, particularly through platforms like OpenAI’s Generative Pre-Trained Transformer (GPT), enables librarians to manage and interpret vast amounts of data, thereby addressing the challenges posed by the data deluge. The paper highlights how AI tools can enhance library services by automating tasks, improving data management, and facilitating the creation of new services, such as chatbots and data collection systems.

The authors propose an experimental framework where data librarians, even without programming experience, can utilize AI-powered tools to generate code in Python based on natural language instructions. This approach aims to enhance the librarians’ algorithmic literacy and hacking skills, enabling them to effectively interact with AI technologies. The study outlines specific applications, including web scraping for bibliometric data from sources like ORCID and Google Scholar, and presents a structured case study that explores the competencies required for data librarians in the AI era. Ultimately, the research underscores the potential of AI to transform library services and empower librarians to leverage technology for innovative data management solutions.

Results

The results section outlines a public initiative aimed at creating scripts and datasets generated by artificial intelligence technologies specifically for non-programming data librarians. This project emphasizes accessibility and usability, ensuring that the resources developed can be effectively utilized by individuals without programming expertise.

All collected data, scripts, and packages are made publicly available in accordance with the Research Data Management Plan (PGD) and are licensed under CC BY 4.0. These resources can be accessed through the Mendeley repository, promoting open science and facilitating broader engagement with AI tools in data management.

Discussion

The discussion on data librarianship emphasizes the evolving role of data librarians, highlighting the necessity for digital, computational, and data literacy in their practice. As data-driven methodologies become integral to academic research, librarians are urged to adapt by acquiring skills related to data management, curation, and visualization. The literature suggests that data literacy is intertwined with information literacy, necessitating a comprehensive understanding of data quality and ethical usage. Furthermore, the emergence of big data has prompted the development of new literacy constructs, including algorithmic literacy, which is essential for navigating AI technologies and understanding the implications of algorithm-driven decision-making.

Algorithmic literacy is defined as the ability to comprehend and apply algorithms in various contexts, which is increasingly vital for data librarians. This literacy encompasses not only programming skills but also the capacity to critically engage with algorithms and their societal impacts. The paper underscores the importance of integrating algorithmic literacy into librarianship education, enabling librarians to leverage AI tools effectively. The authors advocate for a collaborative approach where librarians utilize AI-powered tools, such as OpenAI Codex, to enhance their coding capabilities for tasks like web scraping and data extraction, thereby fostering a more data-literate workforce capable of addressing contemporary challenges in information management.