Implementing Gender Identification from Social Media Text Using Transformer Models: A Practical and Ethical Approach

Authors

  • Mohammed Ali Ibrahim Eltaher Department of Data Science & Artificial Intelligent, Faculty of Information Technology, Tripoli University, Tripoli, Libya المؤلف
  • Ibtusam Abdlsalam Belghassem Department of Data Science & Artificial Intelligent, Faculty of Information Technology, Tripoli University, Tripoli, Libya المؤلف
  • Fatimah Basheer Alqadhi Department of Data Science & Artificial Intelligent, Faculty of Information Technology, Tripoli University, Tripoli, Libya المؤلف
  • Ahmed Alkilany Department of Data Science & Artificial Intelligent, Faculty of Information Technology, Tripoli University, Tripoli, Libya المؤلف

DOI:

https://doi.org/10.35778/169pxc15

Keywords:

Gender identification, transformer models, BERT, social media text, ethical AI, bias mitigation.

Abstract

Gender identification from user-generated text has significant applications in social computing, marketing, and online safety. However, traditional classifiers often fail to capture linguistic subtleties and raise ethical concerns regarding bias and privacy. This study proposes a transformer-based framework for gender classification on social media text, emphasizing transparency and fairness. Using the Facebook/md_gender_bias dataset, we evaluate multiple transformer architectures, including BERT and RoBERTa, comparing their performance against classical machine-learning baselines. The proposed approach achieves improved F1-scores while reducing gender bias through bias-mitigation and explainability techniques. Experimental findings demonstrate the practicality of large language models for gender prediction when combined with ethical constraints and responsible evaluation.

Downloads

Download data is not yet available.

References

[1] P. Mukherjee and J. Liu, “Analysis of gendered communication styles in online forums,” Proc. ACM Conf. Web Sci., 2018, pp. 121–130.

[2] S. Blodgett, S. Barocas, H. Daumé III, and H. Wallach, “Language (technology) is power: A critical survey of ‘bias’ in NLP,” Proc. ACL, 2020, pp. 5454–5476.

[3] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” Proc. NAACL-HLT, 2019, pp. 4171–4186.

[4] Y. Liu et al., “RoBERTa: A robustly optimized BERT pretraining approach,” arXiv preprint arXiv:1907.11692, 2019.

[5] T. Mitchell et al., “Model cards for model reporting,” Proc. FAT, 2019, pp. 220–229.

[6] K. Burger, M. Henderson, and G. Zarrella, “Discriminating gender on Twitter,” Proc. EMNLP Workshop on Author Profiling, 2011, pp. 1301–1309.

[7] A. Argamon, J. Koppel, and M. Fine, “Gender, genre, and writing style in formal written texts,” Text, vol. 23, no. 3, pp. 321–346, 2003.

[8] S. Mukherjee and J. Liu, “Improving gender prediction using stylistic lexical fea-tures,” Proc. ICWSM, 2017, pp. 551–554.

[9] A. K. Jain and R. Ahmad, “Author profiling and gender classification: A compara-tive study,” Inf. Process. Manage., vol. 58, no. 6, pp. 102671, 2021.

[10] T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” Proc. ICLR, 2013.

[11] J. Pennington, R. Socher, and C. Manning, “GloVe: Global vectors for word representation,” Proc. EMNLP, 2014, pp. 1532–1543.

[12] Z. Lan et al., “ALBERT: A Lite BERT for self-supervised learning of language representations,” Proc. ICLR, 2020.

[13] R. Sun and T. Wang, “Evaluating transformer models for gender classification,” IEEE Access, vol. 10, pp. 12035–12047, 2022.

[14] D. Madabushi et al., “Author profiling with BERT: An evaluation on multilin-gual gender prediction,” Proc. CLEF, 2021, pp. 231–245.

[15] M. Yang and S. Lee, “Bias mitigation and interpretability in NLP classification tasks,” IEEE Trans. Affect. Comput., 2023, doi:10.1109/TAFFC.2023.3348294.

[16] J. Chen et al., “Responsible AI for demographic inference,” Proc. AAAI, 2024, pp. 3128–3136.

[17] A. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “1., vol. 55, no. 6, pp. 1–35, 2023.

[18] IEEE, Ethically Aligned Design: A Vision for Prioritizing Human Well-being with Autonomous and Intelligent Systems, 2nd ed., IEEE Standards Assoc., 2021.

[19] K. Crawford and T. Paglen, “Excavating AI: The politics of images in machine learning training sets,” AI & Soc., vol. 37, no. 1, pp. 1–14, 2022.

Downloads

Published

2026-03-31

Journal Volume

Section

المقالات

How to Cite

Implementing Gender Identification from Social Media Text Using Transformer Models: A Practical and Ethical Approach. (2026). Journal of Azzaytuna University, 57, 272-282. https://doi.org/10.35778/169pxc15

Similar Articles

You may also start an advanced similarity search for this article.