The Hidden Bias Of Artificial Intelligence In Indonesian Elt Assessment: A Mixed-Methods Study Of Grammarly And Turnitin Scores On Student Writing

Authors

  • I Putu Andre Suhardiana Universitas Hindu Negeri I Gusti Bagus Sugriwa Denpasar, Indonesia

DOI:

https://doi.org/10.24090/zdy23088

Keywords:

algorithmic bias, automated writing evaluation, decolonial assessment, culturally situated writing, ELT assessment

Abstract

The growing use of AI tools in ELT assessment, like Grammarly and Turnitin, poses risks of algorithmic bias against non-native varieties of English. Although there is extensive research to show bias against Indian, Nigerian and Chinese English writing, no empirical studies about Indonesian English have been investigated nor isolated, comparing AI scores versus human rubric marking which allows culturally situated features like code-switching, politeness markers and more indirect rhetorical structures. This sequential explanatory mixed-methods study examined how Indonesian EFL writing features that are grounded in local culture were penalized by Grammarly and Turnitin, and what lecturers perceived about this bias. Presumably, Grammarly, Turnitin and the three lecturers were invited to score sixteen essays by six university students, then subsequently interviewed semi-structured style, a qualitative design. In the quantitative data, some strong negative bias was revealed: essays were penalized by Grammarly for larger -11.8 points (93.75%) of essays or Turnitin scores -6.9 points (87.5% of essays). The harshest penalties were levied for code-switching (Grammarly bias -15.2), while the moderate ones were reserved for politeness markers (-12.6) and indirect rhetorical structures (-8.9). Lecturers were in agreement against AI-only assessment, devising human-led, AI-assisted models and advocating for intelligibility as the standard reference point instead of native speaker reflective competence. The study concludes that existing AI writing assessment tools will not be appropriate for high-stakes Indonesian ELT assessment without adaptation because they devalue local linguistic identities through systematic ‘cultural erasure,’ as one lecturer called it. 

References

Abdul Rahman, N. A., Zulkornain, L. H., Che Mat, A., & Kustati, M. (2023). Assessing Writing Abilities using AI-Powered Writing Evaluations. Journal of ASIAN Behavioural Studies, 8(24). https://doi.org/10.21834/jabs.v8i24.420

Barrot, J. S. (2023). Using automated written corrective feedback in the writing classrooms: effects on L2 writing accuracy. Computer Assisted Language Learning, 36(4). https://doi.org/10.1080/09588221.2021.1936071

Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2). https://doi.org/10.1191/1478088706qp063oa

Candraningrum, D. (2026). Decolonial Fiction Pedagogy in ELT: Reimagining Memory, History, and Identity Through Laksmi Pamuntjak’s The Question of Red. International Journal of Interdisciplinary Educational Studies, 21(1). https://doi.org/10.18848/2327-011X/CGP/v21i01/49-67

Dikli, S. (2006). An overview of automated scoring of essays. In Journal of Technology, Learning, and Assessment (Vol. 5, Number 1).

Giray, L. (2024). The Problem with False Positives: AI Detection Unfairly Accuses Scholars of AI Plagiarism. In Serials Librarian (Vol. 85, Numbers 5–6). https://doi.org/10.1080/0361526X.2024.2433256

Kaur, A., Noman, M., & Nordin, H. (2017). Inclusive assessment for linguistically diverse learners in higher education. Assessment and Evaluation in Higher Education, 42(5). https://doi.org/10.1080/02602938.2016.1187250

Koltovskaia, S. (2020). Student engagement with automated written corrective feedback (AWCF) provided by Grammarly: A multiple case study. Assessing Writing, 44. https://doi.org/10.1016/j.asw.2020.100450

Manley, S. (2023). The use of text-matching software’s similarity scores. Accountability in Research, 30(4). https://doi.org/10.1080/08989621.2021.1986018

Markl, N. (2022). Language variation and algorithmic bias: understanding algorithmic bias in British English automatic speech recognition. ACM International Conference Proceeding Series. https://doi.org/10.1145/3531146.3533117

Rosa, J., & Flores, N. (2017). Unsettling race and language: Toward a raciolinguistic perspective. Language in Society, 46(5). https://doi.org/10.1017/S0047404517000562

Tajeddin, Z., Saeedi, Z., & Panahzadeh, V. (2022). English Language Teachers’ Perceived Classroom Assessment Knowledge and Practice: Developing and Validating a Scale. Profile: Issues in Teachers’ Professional Development, 24(2). https://doi.org/10.15446/profile.v24n2.90518

Waddington, J. (2024). Questioning the native speaker construct in teacher education: Enabling multilingual identities and decolonial language pedagogies. In Questioning the Native Speaker Construct in Teacher Education: Enabling Multilingual Identities and Decolonial Language Pedagogies. https://doi.org/10.4324/9781003188896

Wahyuningsih, S. (2024). Does Artificial Intelligence (AI) Play Roles in Enhancing Academic Writing? Unravelling Lecturers’ Voices in Indonesian Higher Education. Jurnal Pendidikan Progresif, 14(1). https://doi.org/10.23960/jpp.v14.i1.202436

Weigle, S. C. (2013). English as a second language writing and automated essay evaluation. In Handbook of Automated Essay Evaluation: Current Applications and New Directions. https://doi.org/10.4324/9780203122761

Widodo, H. P. (2018). A Critical Micro-semiotic Analysis of Values Depicted in the Indonesian Ministry of National Education-Endorsed Secondary School English Textbook. In English Language Education (Vol. 9). https://doi.org/10.1007/978-3-319-63677-1_8

Downloads

Published

2026-07-16

How to Cite

The Hidden Bias Of Artificial Intelligence In Indonesian Elt Assessment: A Mixed-Methods Study Of Grammarly And Turnitin Scores On Student Writing. (2026). Conference on English Language Teaching, 6(1), 188-200. https://doi.org/10.24090/zdy23088