DOI: https://doi.org/10.1017/langcog.2026.10090
تاريخ النشر: 2026-01-01
المؤلف: Greg Woodin وآخرون
الموضوع الرئيسي: الجوانب المعرفية والتنموية لمهارات الرياضيات
نظرة عامة
في هذا القسم، يستكشف المؤلفون استخدام الأرقام المستديرة في اللغة، مع تسليط الضوء على دقتها الرياضية مقارنة باستخدامها التقريبي في السياقات اليومية. من خلال ثلاث تحليلات للغة الإنجليزية البريطانية والأمريكية، يجدون أن الأفراد يميلون إلى استخدام الأرقام المستديرة بشكل أكثر تكرارًا والتقريب إلى حد أكبر مع زيادة الحجم. ومن الجدير بالذكر أن دلالات الأرقام المستديرة الأكبر تظهر تشابهات مع الأرقام اللامحددة المفرطة، مثل “غازيلون”، التي لا تنقل قيمًا دقيقة.
بالإضافة إلى ذلك، يلاحظ المؤلفون أن الألغاز الكبيرة غالبًا ما تظهر اختلافات كبيرة بين عدد القطع المعلن وعدد القطع الفعلي، مما يوضح هذه النزعة (على سبيل المثال، الإعلان عن “13,200 قطعة” عندما يكون العدد الفعلي 13,224). يجادلون بأن هذه الظاهرة متجذرة في بنية الأرقام الإنجليزية القائمة على الرقم 10، والتي تؤكد على قوى 10 وتسهّل التقريب في التواصل. يتماشى هذا مع نظام الأعداد التقريبية، حيث يتم تمثيل الكميات الأكبر بشكل معرفي مع زيادة عدم الدقة، مما يعقد التقدير والمقارنة.
مقدمة
في مقدمة هذه الورقة البحثية، يستكشف المؤلفون الاستخدام الدقيق للتعبيرات العددية، وخاصة التمييز بين الأرقام المستديرة وغير المستديرة في اللغة اليومية. يوضحون أن مصطلحات مثل “مئة” غالبًا ما تعمل كتقديرات بدلاً من أعداد دقيقة، وهي ظاهرة وصفها لاسرسون (1999) بأنها “هالة براغماتية”. يتناقض هذا مع الأرقام المحددة، مثل “سبعة وتسعين”، التي تشير إلى الدقة. يناقش المؤلفون أيضًا كيف يؤثر حجم الأرقام على استخدامها، مشيرين إلى أن الأرقام المستديرة الأكبر (مثل “مليون” أو “مليار”) من المرجح أن تُستخدم كعلامات تقريبية بدلاً من أرقام دقيقة، خاصة في السياقات المفرطة.
تقدم الورقة تحليلات كمية تُظهر أن الأرقام المستديرة تُستخدم بشكل أكثر تكرارًا وتُقرب إلى حد أكبر عند الأحجام الأعلى. كما تكشف أن الأرقام المستديرة الأكبر تشترك في سياقات لغوية مع الأرقام المفرطة غير المحددة، مثل “غازيلون”. يقترح المؤلفون أن هذه العلاقة تنبع من بنية الأرقام الإنجليزية القائمة على الرقم 10، حيث تشغل بعض الأرقام نقاط بارزة على خط الأعداد، مما يسهل استخدامها للتقريب. يربطون هذه النتائج بنظام الأعداد التقريبية (ANS)، مقترحين أن الأرقام الأكبر تُمثل معرفيًا بطريقة أكثر تقريبًا من الأرقام الأصغر، مما يتماشى مع الأدبيات الموجودة حول الإدراك العددي.
طرق البحث
في هذه الدراسة، استخدم المؤلفون تمثيلات الكلمات ‘word2vec’، المدربة مسبقًا على مجموعة بيانات ويكيبيديا، لتحليل كلمات الأرقام الإنجليزية، بما في ذلك الأعداد الصحيحة (1-99)، وقوى العشرة، والأرقام المفرطة غير المحددة. تم اختيار word2vec بناءً على سهولة الوصول إليه وأدائه المتفوق مقارنةً بالبدائل الأخرى في تطبيقات مختلفة. ركز التحليل على الكلمات الفردية والأرقام الموصولة، مستبعدًا التعبيرات متعددة الكلمات، لضمان قابلية المقارنة مع الأرقام المفرطة التي تفتقر إلى تمثيلات عددية. تم تمثيل التضمينات في فضاء متجه مستمر بعدد أبعاد 300، حيث يتم وضع الكلمات ذات الصلة دلاليًا بالقرب من بعضها البعض.
لقياس التشابهات الدلالية بين كلمات الأرقام، قام المؤلفون بحساب تشابهات جيب التمام عبر الأبعاد الـ 300. تشير هذه المقياس إلى مدى استخدام أزواج الأرقام في سياقات مشابهة. للتصور، تم استخدام التحجيم متعدد الأبعاد (MDS) لتقليل الأبعاد من 300 إلى 2، مع محاولة الحفاظ على التشابهات الكامنة. يضمن خوارزمية MDS أن الكلمات القريبة في أبعاد متعددة تبقى كذلك في الفضاء المخفض، حتى لو كانت بعيدة في بعد واحد، مما يوفر تمثيلًا ذا معنى للعلاقات بين كلمات الأرقام.
النتائج
يقدم قسم النتائج النتائج الرئيسية المتعلقة بالتفاعل بين خصائص الاستدارة وLog 10 Magnitude، كما هو موضح في الشكل 1. جميع معاملات التفاعل إيجابية، مما يشير إلى أن الأرقام المستديرة تُستخدم بشكل أكثر تكرارًا عند الأحجام الأعلى في كل من مجموعة البيانات الوطنية البريطانية (BNC) ومجموعة بيانات اللغة الإنجليزية الأمريكية المعاصرة (COCA). ومن الجدير بالذكر أن تقدير الميل لـ “مضاعف 5” يظهر عدم يقين كبير، حيث أن فترة الثقة 95% الخاصة به تقع بشكل هامشي فوق 0 في BNC وتحتوي على 0 في COCA. بالمقابل، تظهر مصطلحات التفاعل الأخرى ميلًا إيجابيًا قويًا، مع فترات ثقة أعلى بكثير من 0، مرتبة حسب القوة المتزايدة كما يلي: مضاعف 10، 2-ness، 5-ness، 2.5-ness، و10-ness. يكشف تحليل الترددات الخام عن انخفاض في الترددات المطلقة مع زيادة الحجم، كما يتضح من تردد BNC للعدد 100 (14,819) مقارنةً بـ 1,000 (6,975).
يوضح الشكل 2 تمثيلات الكلمات في فضاء التحجيم متعدد الأبعاد (MDS) ثنائي الأبعاد، كاشفًا عن أنماط تجميع مميزة. تتجمع الأرقام الأصغر (2 إلى 99) على اليمين، بينما تتواجد الأرقام الأكبر في الأعلى، مما يشير إلى توزيع يعتمد على الحجم. تشكل الأرقام المفرطة غير المحددة وقوى العشرة الأكبر (مثل “مليون”، “مليار”) مجموعة منفصلة على اليسار، مما يشير إلى استخدامها بشكل غير دقيق، على غرار الأرقام المفرطة غير المحددة. تشغل قوى العشرة الأصغر (“مئة” و”ألف”) مساحة انتقالية، مما يشير إلى احتمال معتدل للاستخدام غير الدقيق. قد يمثل المحور الأفقي لـ MDS استمرارية احتمال عدم الدقة، حيث يزداد من اليمين إلى اليسار. بالإضافة إلى ذلك، تحتوي أقل من نصف الألغاز التي تم تحليلها (44.1%) على العدد الدقيق من القطع المعلن عنها، مع اختلافات أكثر وضوحًا في الألغاز الأكبر (≥1,000 قطعة) مقارنةً بالألغاز الأصغر (<1,000 قطعة)، كما هو موضح في الشكل 3 ومُلخص في الجدول 5.
مناقشة
يتناول قسم المناقشة في الورقة البحثية مفهوم الأرقام المستديرة، مؤكدًا على أهميتها الهيكلية والنفسية ضمن الأنظمة العددية. تُعرف الأرقام المستديرة بأنها تلك التي تكون بارزة هيكليًا بسبب علاقتها بالأساس وقواه، وتظهر “استدارة” متدرجة بناءً على عدد هذه العلاقات. على سبيل المثال، بينما يُعتبر 310 مستديرًا كمضاعف لـ 10، يُعتبر 500 أكثر استدارة بسبب خصائصه الإضافية كمضاعف لكل من 10 و100، بالإضافة إلى كونه نصف 1,000. يتم عكس هذه التميز الهيكلي في اللغة، حيث تكون الأرقام المستديرة مثل “ألف” أبسط وأكثر بروزًا من نظيراتها غير المستديرة.
تستكشف الورقة أيضًا وظائف الأرقام المستديرة، مصنفة إياها إلى استخدامات تقريبية وغير تقريبية. تشمل الوظائف التقريبية تقريب القيم من أجل الكفاءة التواصلية أو التعبير عن عدم اليقين، بينما تتضمن الوظائف غير التقريبية استخدام الأرقام المستديرة كمعايير أو مراجع ثقافية. يشير المؤلفون إلى أن انتشار الأرقام المستديرة في اللغة والبيانات يعكس بروزها النفسي، مما يخلق حلقة تغذية راجعة تعزز استخدامها المتكرر. بالإضافة إلى ذلك، يتم تسليط الضوء على العلاقة بين الاستدارة والحجم، مع وجود نتائج تشير إلى أن الأرقام المستديرة تُستخدم بشكل أكثر تكرارًا عند الأحجام الأعلى، مما يشير إلى ميل نحو التقريب. تتماشى هذه الاتجاهات عبر مجموعات بيانات مختلفة، مما يعزز فكرة أن الأرقام المستديرة تعمل كأدوات معرفية وتواصلية مهمة في التعبير العددي.
DOI: https://doi.org/10.1017/langcog.2026.10090
Publication Date: 2026-01-01
Author(s): Greg Woodin et al.
Primary Topic: Cognitive and developmental aspects of mathematical skills
Overview
In this section, the authors explore the use of round numbers in language, highlighting their mathematical precision contrasted with their approximate usage in everyday contexts. Through three analyses of British and American English, they find that individuals tend to use round numbers more frequently and round to a greater extent as the magnitude increases. Notably, the semantics of larger round numbers exhibit similarities to indefinite hyperbolic numbers, such as “gazillion,” which do not convey precise values.
Additionally, the authors observe that larger jigsaw puzzles often show significant discrepancies between advertised piece counts and actual counts, exemplifying this tendency (e.g., advertising “13,200 pieces” when the actual count is 13,224). They argue that this phenomenon is rooted in the base-10 structure of English numerals, which emphasizes powers of 10 and facilitates approximation in communication. This aligns with the approximate number system, where larger quantities are cognitively represented with increasing imprecision, complicating estimation and comparison.
Introduction
In the introduction of this research paper, the authors explore the nuanced use of numerical expressions, particularly the distinction between round and non-round numbers in everyday language. They illustrate that terms like “a hundred” often serve as estimates rather than precise counts, a phenomenon described by Lasersohn (1999) as a “pragmatic halo.” This contrasts with specific numbers, such as “ninety-seven,” which imply exactness. The authors further discuss how the magnitude of numbers influences their usage, noting that larger round numbers (e.g., “million” or “billion”) are more likely to be employed as approximate markers rather than exact figures, especially in hyperbolic contexts.
The paper presents quantitative analyses demonstrating that round numbers are used more frequently and rounded to a greater extent at higher magnitudes. It also reveals that larger round numbers share linguistic contexts with indefinite hyperbolic numerals, such as “gazillion.” The authors propose that this relationship stems from the base-10 structure of English numerals, where certain numbers occupy salient points on the number line, facilitating their use for approximation. They connect these findings to the approximate number system (ANS), suggesting that larger numbers are cognitively represented in a more approximate manner than smaller ones, aligning with existing literature on numerical cognition.
Methods
In this study, the authors employed ‘word2vec’ word embeddings, pre-trained on the Wikipedia corpus, to analyze English number words, including integers (1-99), powers of ten, and indefinite hyperbolic numbers. The choice of word2vec was based on its accessibility and superior performance compared to other alternatives in various applications. The analysis focused on single words and hyphenated numbers, excluding multi-word expressions, to ensure comparability with hyperbolic numbers that lack numeral representations. The embeddings were represented in a continuous 300-dimensional vector space, where semantically related words are positioned closer together.
To quantify the semantic similarities among the number words, the authors calculated cosine similarities across the 300 dimensions. This metric indicates the extent to which pairs of numbers are used in similar contexts. For visualization, multidimensional scaling (MDS) was utilized to reduce the dimensionality from 300 to 2, while attempting to preserve the inherent similarities. The MDS algorithm ensures that words that are close in multiple dimensions remain so in the reduced space, even if they are distant in one dimension, thereby providing a meaningful representation of the relationships among the number words.
Results
The results section presents key findings regarding the interaction between roundness properties and Log 10 Magnitude, as illustrated in Figure 1. All interaction coefficients are positive, indicating that round numbers are more frequently utilized at higher magnitudes in both the British National Corpus (BNC) and the Corpus of Contemporary American English (COCA). Notably, the slope estimate for “Multiple of 5” shows considerable uncertainty, with its 95% credible interval marginally above 0 in the BNC and including 0 in COCA. In contrast, other interaction terms exhibit robust positive slopes, with credible intervals well above 0, ordered by increasing strength as follows: Multiple of 10, 2-ness, 5-ness, 2.5-ness, and 10-ness. The analysis of raw frequencies reveals a decline in absolute frequencies with increasing magnitude, exemplified by the BNC’s frequency of 100 (14,819) compared to 1,000 (6,975).
Figure 2 depicts word embeddings in a two-dimensional multidimensional scaling (MDS) space, revealing distinct clustering patterns. Smaller numbers (2 to 99) cluster on the right, with larger numbers positioned higher, indicating a magnitude-dependent distribution. Indefinite hyperbolic numbers and larger powers of 10 (e.g., ‘million’, ‘billion’) form a separate cluster on the left, suggesting they are used imprecisely, akin to indefinite hyperbolic numbers. Smaller powers of 10 (‘hundred’ and ‘thousand’) occupy a transitional space, indicating a moderate likelihood of imprecise usage. The MDS horizontal axis may represent a continuum of imprecision probability, increasing from right to left. Additionally, fewer than half of the jigsaw puzzles analyzed (44.1%) contained the exact number of pieces advertised, with discrepancies more pronounced in larger puzzles (≥1,000 pieces) compared to smaller ones (<1,000 pieces), as shown in Figure 3 and summarized in Table 5.
Discussion
The discussion section of the research paper elaborates on the concept of round numbers, emphasizing their structural and psychological significance within numerical systems. Round numbers, defined as those that are structurally salient due to their relation to the base and its powers, exhibit graded ’roundness’ based on the number of such relations. For instance, while 310 is considered round as a multiple of 10, 500 is deemed rounder due to its additional properties of being a multiple of both 10 and 100, as well as being half of 1,000. This structural distinctiveness is mirrored in language, where round numbers like ‘thousand’ are simpler and more salient than their non-round counterparts.
The paper further explores the functions of round numbers, categorizing them into approximate and non-approximate uses. Approximate functions include rounding values for communicative efficiency or expressing uncertainty, while non-approximate functions involve using round numbers as benchmarks or cultural references. The authors note that the prevalence of round numbers in language and data reflects their psychological salience, creating a feedback loop that reinforces their frequent use. Additionally, the relationship between roundness and magnitude is highlighted, with findings indicating that round numbers are more frequently used at higher magnitudes, suggesting a tendency towards approximation. This trend is consistent across different corpora, reinforcing the idea that round numbers serve as important cognitive and communicative tools in numerical expression.
