On the Representation of Racial and Ethnic Subgroups in AI-generated Texts: A Case Study in Automated Essay Scoring
DOI:
https://doi.org/10.64634/ac01td58Keywords:
artificial intelligence in education, automated essay scoring, bias, large language models, LLMs, AI, artificial intelligence, natural language processing, NLPAbstract
In this study, we assess the capability of large language models (LLMs) to generate essays by a specific subgroup (i.e., race/ethnicity) after being given example essays and rubrics; we then investigate the efficacy of data augmented in this manner for automated essay scoring with respect to model performance and bias. In a series of experiments, we use models GPT-4 and GPT-4o and ask them to generate essays from a given subgroup after inferring the race/ethnicity of the writer. We find that while LLMs can be directed to generate essays for specific demographic groups, the inferred racial and ethnic distribution in the generated data does not closely mirror the actual distribution observed in the source data set. We augment existing data for underrepresented subgroups with LLM-generated data separated into two groups—with correct LLM race prediction, and with incorrect race prediction—and assess the improvement in agreement with human scores with quadratic weighted kappa and bias mitigation as change in standardized mean difference. Our analysis shows that while LLMs struggle to predict the race accurately from given samples, augmentation with such data can be helpful to mitigate bias regardless.
Suggested citation: Badola, A., Zhang, M., & Li, Chen. (2026). On the representation of racial and ethnic subgroups in AI-generated texts: A case study in automated essay scoring (Research Memorandum No. RM–26-06). ETS. https://doi.org/10.64634/ac01td58
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Educational Testing Service

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.