Implementing Gender Identification from Social Media Text Using Transformer Models: A Practical and Ethical Approach

المؤلفون

  • Mohammed Eltaher Department of Data Science & Artificial Intelligent, Faculty of Information Technology, Tripoli University, Tripoli, Libya المؤلف
  • Ibtusam Belghassem Department of Data Science & Artificial Intelligent, Faculty of Information Technology, Tripoli University, Tripoli, Libya المؤلف
  • Fatimah Alqadhi Department of Data Science & Artificial Intelligent, Faculty of Information Technology, Tripoli University, Tripoli, Libya المؤلف
  • Ahmed Alkilany Department of Data Science & Artificial Intelligent, Faculty of Information Technology, Tripoli University, Tripoli, Libya المؤلف

DOI:

https://doi.org/10.35778/169pxc15

الكلمات المفتاحية:

تحديد الجنس، نماذج المحولات، BERT، نصوص وسائل التواصل الاجتماعي، الذكاء الاصطناعي الأخلاقي، تخفيف التحيز

الملخص

يُعدّ تحديد الجنس من النصوص التي يُنشئها المستخدمون ذا أهمية بالغة في مجالات الحوسبة الاجتماعية والتسويق والسلامة على الإنترنت. مع ذلك، غالبًا ما تعجز المصنفات التقليدية عن استيعاب الفروق اللغوية الدقيقة، وتثير مخاوف أخلاقية تتعلق بالتحيز والخصوصية. تقترح هذه الدراسة إطار عمل قائم على نموذج Transformer لتصنيف الجنس في نصوص وسائل التواصل الاجتماعي، مع التركيز على الشفافية والإنصاف. باستخدام مجموعة بيانات facebook/md_gender_bias، نقوم بتقييم عدة نماذج Transformer، بما في ذلك BERT وRoBERTa، ومقارنة أدائها مع نماذج التعلم الآلي التقليدية. يحقق النهج المقترح تحسينًا في مقياس F1 مع تقليل التحيز الجنسي من خلال تقنيات تخفيف التحيز وتفسير النتائج. تُظهر النتائج التجريبية جدوى استخدام نماذج اللغة الكبيرة للتنبؤ بالجنس عند دمجها مع القيود الأخلاقية والتقييم المسؤول.

التنزيلات

تنزيل البيانات ليس متاحًا بعد.

المراجع

[1] P. Mukherjee and J. Liu, “Analysis of gendered communication styles in online forums,” Proc. ACM Conf. Web Sci., 2018, pp. 121–130.

[2] S. Blodgett, S. Barocas, H. Daumé III, and H. Wallach, “Language (technology) is power: A critical survey of ‘bias’ in NLP,” Proc. ACL, 2020, pp. 5454–5476.

[3] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” Proc. NAACL-HLT, 2019, pp. 4171–4186.

[4] Y. Liu et al., “RoBERTa: A robustly optimized BERT pretraining approach,” arXiv preprint arXiv:1907.11692, 2019.

[5] T. Mitchell et al., “Model cards for model reporting,” Proc. FAT, 2019, pp. 220–229.

[6] K. Burger, M. Henderson, and G. Zarrella, “Discriminating gender on Twitter,” Proc. EMNLP Workshop on Author Profiling, 2011, pp. 1301–1309.

[7] A. Argamon, J. Koppel, and M. Fine, “Gender, genre, and writing style in formal written texts,” Text, vol. 23, no. 3, pp. 321–346, 2003.

[8] S. Mukherjee and J. Liu, “Improving gender prediction using stylistic lexical fea-tures,” Proc. ICWSM, 2017, pp. 551–554.

[9] A. K. Jain and R. Ahmad, “Author profiling and gender classification: A compara-tive study,” Inf. Process. Manage., vol. 58, no. 6, pp. 102671, 2021.

[10] T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” Proc. ICLR, 2013.

[11] J. Pennington, R. Socher, and C. Manning, “GloVe: Global vectors for word representation,” Proc. EMNLP, 2014, pp. 1532–1543.

[12] Z. Lan et al., “ALBERT: A Lite BERT for self-supervised learning of language representations,” Proc. ICLR, 2020.

[13] R. Sun and T. Wang, “Evaluating transformer models for gender classification,” IEEE Access, vol. 10, pp. 12035–12047, 2022.

[14] D. Madabushi et al., “Author profiling with BERT: An evaluation on multilin-gual gender prediction,” Proc. CLEF, 2021, pp. 231–245.

[15] M. Yang and S. Lee, “Bias mitigation and interpretability in NLP classification tasks,” IEEE Trans. Affect. Comput., 2023, doi:10.1109/TAFFC.2023.3348294.

[16] J. Chen et al., “Responsible AI for demographic inference,” Proc. AAAI, 2024, pp. 3128–3136.

[17] A. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “1., vol. 55, no. 6, pp. 1–35, 2023.

[18] IEEE, Ethically Aligned Design: A Vision for Prioritizing Human Well-being with Autonomous and Intelligent Systems, 2nd ed., IEEE Standards Assoc., 2021.

[19] K. Crawford and T. Paglen, “Excavating AI: The politics of images in machine learning training sets,” AI & Soc., vol. 37, no. 1, pp. 1–14, 2022.

التنزيلات

منشور

2026-03-31

إصدار

القسم

المقالات

كيفية الاقتباس

Implementing Gender Identification from Social Media Text Using Transformer Models: A Practical and Ethical Approach. (2026). مجلة جامعة الزيتونة , 57, 272-282. https://doi.org/10.35778/169pxc15

المؤلفات المشابهة

يمكنك أيضاً إبدأ بحثاً متقدماً عن المشابهات لهذا المؤلَّف.