DOI: https://doi.org/10.1038/s41598-025-17831-3
PMID: https://pubmed.ncbi.nlm.nih.gov/41044255
تاريخ النشر: 2025-10-03
المؤلف: John Valerian Corda وآخرون
الموضوع الرئيسي: دراسات الأنثروبولوجيا الجنائية والبيوآركيولوجيا
نظرة عامة
تدرس هذه الدراسة تطبيق خوارزميات التعلم الآلي لتقدير الاختلاف الجنسي بناءً على الميزات غير المترية للفك السفلي، وهو جانب حاسم في علم الأنثروبولوجيا الجنائية. باستخدام مجموعة بيانات تضم 156 فردًا من كارناتاكا، الهند، تقيم البحث أربعة خوارزميات: الجيران الأقرب (KNN)، شجرة القرار (DT)، آلات الدعم الناقل (SVM)، والغابة العشوائية (RF)، مع التركيز على 12 معلمة غير مترية للفك السفلي. تشير النتائج إلى أن خوارزمية الغابة العشوائية تفوقت على الآخرين، محققة أعلى مؤشر جاكارد (0.86)، ودرجة F1 (0.92)، ودقة (0.92) عبر كل من طرق SMOTE وRandom Over-Sampling (ROS). ومن الجدير بالذكر أن زاوية N6 غونال وحدود N12 انحناء رامال الخلفية تم تحديدها كمؤشرات هامة للجنس.
في الختام، كشفت التحليل المقارن للخوارزميات الأربعة أنه بينما أظهرت جميع النماذج قدرات تصنيفية متقاربة، فإن الغابة العشوائية أظهرت باستمرار مقاييس أداء متفوقة، بما في ذلك الدقة والقدرة على مقاومة عدم توازن البيانات. كما أبرزت الدراسة أن ROS حسنت أداء KNN وDT وSVM، على الرغم من أن هذه النماذج لم تصل إلى القوة التنبؤية للغابة العشوائية. تؤكد النتائج فعالية التعلم الآلي في التوصيف البيولوجي وتوصي باستخدام الغابة العشوائية كخيار مثالي لتصنيف الجنس بناءً على ميزات الفك السفلي، نظرًا لاستقراره وقدرته على إدارة تفاعلات الميزات المعقدة.
طرق
في هذه الدراسة، تم استخدام نهج متعدد التخصصات يجمع بين الطب والهندسة لتحليل ميزات الفك السفلي لتصنيف الجنس. تضمنت البحث توثيقًا منهجيًا وتصنيفًا للبيانات النوعية غير المترية من مجموعة بيانات تضم 102 فك سفلي ذكري و54 فك سفلي أنثوي، مما يبرز عدم توازن كبير في الفئات قد يؤثر على نماذج التعلم الآلي (ML). لمعالجة هذه المشكلة، استخدم الباحثون تقنية الزيادة الاصطناعية للأقليات (SMOTE) وRandom Over-sampling (ROS) لتحقيق توازن بين الفئات، مما يعزز دقة النموذج للمجموعة الأقلية.
كما شملت المنهجية تطبيق ترميز one-hot لتحويل الميزات الفئوية إلى قيم ثنائية، مما يضمن أن ميزات الفك السفلي غير المترية تم التعامل معها كمتغيرات مستقلة. منع هذا النهج الافتراضات العلائقية المضللة وحسن من قابلية تفسير الميزات ضمن خوارزميات تصنيف ML. تم تقديم أوصاف مفصلة لطرق جمع البيانات والخوارزميات المحددة المستخدمة في الأقسام التالية من الورقة.
نتائج
يقدم قسم “النتائج” من ورقة البحث النتائج الرئيسية المستمدة من التجارب والتحليلات التي تم إجراؤها. يبرز النتائج المهمة التي تدعم الفرضيات أو أسئلة البحث المطروحة سابقًا في الدراسة. يتم تمثيل البيانات عادةً من خلال أشكال متنوعة، بما في ذلك الجداول والرسوم البيانية والتحليلات الإحصائية، التي توضح العلاقات والاتجاهات الملاحظة في مجموعة البيانات.
تشير النتائج إلى أن المتغيرات الرئيسية ذات الاهتمام تظهر ارتباطًا ذا دلالة إحصائية، كما يتضح من قيم p التي تقل عن 0.05. بالإضافة إلى ذلك، تكشف التحليلات أن النموذج المستخدم للتنبؤات يظهر درجة عالية من الدقة، مع قيمة R-squared تتجاوز 0.85. تشير هذه النتائج إلى أن الإطار النظري المقترح قوي وقابل للتطبيق على الظواهر المدروسة، مما يوفر أساسًا لمزيد من البحث والتطبيقات العملية المحتملة.
مناقشة
ت outlines قسم المناقشة في هذه الورقة البحثية المنهجية والنتائج لدراسة رصدية متعددة المراكز تشمل 156 فك سفلي إنساني بالغ من جنوب الهند. ركزت الدراسة على 12 معلمة غير مترية تتعلق بشكل الفك السفلي، مع ملاحظات قام بها خمسة مؤلفين ذوي خبرة لضمان الموثوقية. تم استخدام التحليلات الإحصائية، بما في ذلك اختبارات كاي-تربيع وإحصائيات كابا، لتقييم تباين الملاحظات بين المراقبين، مما يكشف عن اختلافات كبيرة في بعض المعلمات بين الفك السفلي الذكري والأنثوي، مثل شكل الحدود السفلية وزاوية الغونال.
فيما يتعلق بنماذج التعلم الآلي (ML)، قامت الدراسة بتقييم عدة خوارزميات، بما في ذلك الجيران الأقرب (KNN)، أشجار القرار (DT)، آلات الدعم الناقل (SVM)، والغابة العشوائية (RF)، لتصنيف الفكوك بناءً على المعلمات الملاحظة. أظهر نموذج RF أعلى دقة ودرجة F1، مما يدل على قوته في التعامل مع عدم توازن الفئات من خلال تقنيات مثل تقنية الزيادة الاصطناعية للأقليات (SMOTE) وRandom Oversampling (ROS). سلطت النتائج الضوء على أنه بينما كان أداء KNN الأقل فعالية، قدم كل من SVM وDT توازنًا جيدًا بين الدقة والاسترجاع. بشكل عام، تؤكد النتائج فعالية طرق التجميع مثل RF في تحقيق نتائج تصنيف موثوقة، خاصة في مجموعات البيانات التي تتميز بتوزيعات غير متوازنة.
القيود
تقدم الدراسة عدة قيود قد تؤثر على عمومية نتائجها. أولاً، كانت حجم العينة للمعلمات غير المترية صغيرًا نسبيًا، مما قد يحد من قوة النتائج. بالإضافة إلى ذلك، كان هناك عدم توازن ملحوظ في الجنس في العينة، مع 102 مشارك ذكري و54 مشارك أنثوي. تم معالجة هذا التوازن باستخدام تقنيات الزيادة مثل SMOTE (تقنية الزيادة الاصطناعية للأقليات) وROS (الزيادة العشوائية). علاوة على ذلك، تم جمع البيانات حصريًا من منطقة جغرافية واحدة، وهي سكان جنوب الهند، مما قد لا يمثل تنوعات ديموغرافية أوسع.
لمعالجة هذه القيود، ستركز الأبحاث المستقبلية على زيادة حجم العينة وتنويع التمثيل الجغرافي للمشاركين. يهدف هذا النهج إلى تعزيز موثوقية النتائج واستكشاف الفروق المحتملة عبر مختلف السكان.
DOI: https://doi.org/10.1038/s41598-025-17831-3
PMID: https://pubmed.ncbi.nlm.nih.gov/41044255
Publication Date: 2025-10-03
Author(s): John Valerian Corda et al.
Primary Topic: Forensic Anthropology and Bioarchaeology Studies
Overview
The study investigates the application of machine learning algorithms for estimating sexual dimorphism based on non-metric features of the mandible, a critical aspect of forensic anthropology. Utilizing a dataset of 156 individuals from Karnataka, India, the research evaluates four algorithms: k-nearest neighbors (KNN), decision tree (DT), support vector machines (SVM), and random forest (RF), focusing on 12 mandibular non-metric parameters. The results indicate that the Random Forest algorithm outperformed the others, achieving the highest Jaccard Index (0.86), F1 score (0.92), and accuracy (0.92) across both SMOTE and Random Over-Sampling (ROS) methods. Notably, the N6 Gonial angle and N12 Flexure ramal post border were identified as significant predictors of sex.
In conclusion, the comparative analysis of the four algorithms revealed that while all models demonstrated comparable classification capabilities, Random Forest consistently exhibited superior performance metrics, including accuracy and robustness against data imbalance. The study also highlighted that ROS improved the performance of KNN, DT, and SVM, although these models did not reach the predictive power of Random Forest. The findings underscore the effectiveness of machine learning in biological profiling and recommend Random Forest as the optimal choice for sex classification based on mandibular features, given its stability and ability to manage complex feature interactions.
Methods
In this study, a multidisciplinary approach integrating medicine and engineering was employed to analyze mandible features for gender classification. The research involved the systematic documentation and categorization of qualitative non-metric data from a dataset comprising 102 male and 54 female mandibles, highlighting a significant class imbalance that could bias machine learning (ML) models. To address this issue, the researchers utilized the Synthetic Minority Over-sampling Technique (SMOTE) and Random Over-sampling (ROS) to balance the classes, thereby enhancing the model’s accuracy for the minority group.
The methodology also included the application of one-hot encoding to convert categorical features into binary values, ensuring that the non-metric mandibular features were treated as independent variables. This approach prevented misleading relational assumptions and improved the interpretability of the features within the ML classification algorithms. Detailed descriptions of the data collection methods and the specific ML algorithms employed are provided in subsequent sections of the paper.
Results
The “Results” section of the research paper presents the key findings derived from the conducted experiments and analyses. It highlights the significant outcomes that support the hypotheses or research questions posed earlier in the study. The data is typically represented through various forms, including tables, graphs, and statistical analyses, which illustrate the relationships and trends observed in the dataset.
The results indicate that the primary variables of interest exhibit a statistically significant correlation, as evidenced by p-values less than 0.05. Additionally, the analysis reveals that the model used for predictions demonstrates a high degree of accuracy, with an R-squared value exceeding 0.85. These findings suggest that the proposed theoretical framework is robust and applicable to the studied phenomena, providing a foundation for further research and potential practical applications.
Discussion
The discussion section of this research paper outlines the methodology and findings of a multicenter cross-sectional observational study involving 156 adult human mandibles from South India. The study focused on 12 non-metric parameters related to mandibular morphology, with observations made by five experienced authors to ensure reliability. Statistical analyses, including Chi-square tests and Kappa statistics, were employed to assess intra- and inter-observer variability, revealing significant differences in certain parameters between male and female mandibles, such as the shape of the lower border and the gonial angle.
In terms of machine learning (ML) models, the study evaluated several algorithms, including K-Nearest Neighbors (KNN), Decision Trees (DT), Support Vector Machines (SVM), and Random Forest (RF), to classify the mandibles based on the observed parameters. The RF model demonstrated the highest accuracy and F1 score, indicating its robustness in handling class imbalance through techniques like Synthetic Minority Oversampling Technique (SMOTE) and Random Oversampling (ROS). The results highlighted that while KNN performed the least effectively, both SVM and DT provided a good balance of precision and recall. Overall, the findings underscore the efficacy of ensemble methods like RF in achieving reliable classification outcomes, particularly in datasets characterized by imbalanced distributions.
Limitations
The study presents several limitations that may affect the generalizability of its findings. Firstly, the sample size for nonmetric parameters was relatively small, which could limit the robustness of the results. Additionally, there was a notable gender imbalance in the sample, with 102 male and 54 female participants. This imbalance was addressed using oversampling techniques such as SMOTE (Synthetic Minority Over-sampling Technique) and ROS (Random Over-Sampling). Furthermore, the data were exclusively collected from a single geographical region, specifically the South Indian population, which may not represent broader demographic variations.
To address these limitations, future research will focus on increasing the sample size and diversifying the geographical representation of participants. This approach aims to enhance the reliability of the findings and to explore potential differences across various populations.
