Researchers at Princeton University and the University of Chicago found that large language models (LLMs), including tools like ChatGPT, can develop biases not only from human training data but also from their own experiences in simulated hiring scenarios. In these experiments, models segregated candidates based on demographic information after interpreting early outcomes, showing a heightened tendency to stereotype job applicants compared to human participants.
The study utilized a simulated hiring game involving 20 job roles, such as doctors and janitors, across four fictional ethnic groups: Tufa, Aima, Reku, and Weki. Each model was tasked with maximizing successful hires over 40 rounds, without knowing that all candidates had equal chances of success. Models frequently adapted their hiring preferences after learning that specific candidates failed or succeeded in roles. For instance, if an Aima candidate did not perform well as a doctor, models shifted to hiring Aimas for janitorial positions instead.
According to the research, LLMs scored significantly higher on a segregation scale than human subjects, averaging a score of approximately 1.83, compared to 0.84 for humans. Coauthor Ryan Liu, a PhD student at Princeton, stated, “LLMs really are eager to create generalizations from limited data.” This tendency to generalize quickly can lead to damaging stereotypes, as newer models demonstrated even stronger biases.
The implications are significant, especially as AI chatbots gain advanced memory and personalization features. Angelina Wang, a computer scientist at Cornell University, remarked that chatbots could form entrenched biases by over-relying on previous interactions. Simply reducing their memory is not a solution, as users prefer systems that remember their inputs.
The study revealed that instructing the models to prioritize fairness showed limited impact on their behavior. However, incentivizing diverse hiring resulted in less biased decision-making. Liu noted that integrating social goals into AI objectives could promote more socially responsible behaviors in language models.
Additionally, LLMs displayed reduced biases when given more relevant personal information about candidates, such as age and education, as opposed to irrelevant details like hair color. The study highlights the uncertainty surrounding the real-world application of these findings, particularly since AI models screening job applications do not receive immediate feedback on hires.
The researchers raised concerns that as companies increasingly deploy LLMs for résumé screening and other critical evaluations, there exists a risk of the models developing novel biases without direct human input. “These novel biases—they’re sort of ever present,” Liu said, emphasizing the need for awareness as AI use in hiring and other decision-making processes expands.





