
AI is more likely than humans to form biases when hiring
Quick Answer
New research indicates that LLMs like OpenAI's o3 and Claude exhibit stronger biases in hiring than humans, with a segregation score of 1.83 compared to 0.84 for human participants.
Quick Take
These biases arise from limited data generalizations and can be mitigated by incorporating social values into model objectives.
Key Points
- developed biases from hiring simulations, stereotyping candidates by demographic group.
- OpenAI's o3 scored 1.83 on the segregation scale, significantly higher than human scores.
- Models became less biased when given personal information about candidates.
- Incorporating social values into model objectives reduced bias in hiring decisions.
- AI's tendency to generalize from limited data can lead to novel biases not taught by humans.
DeepSignal Analysis
What happened
Recent research indicates that large language models (LLMs) like OpenAI's o3 and Anthropic's Claude exhibit stronger biases in hiring compared to human participants. In a simulated hiring game, LLMs scored a segregation score of 1.83, significantly higher than the human score of 0.84. These biases stem from the models' tendency to generalize from limited data and can be mitigated by incorporating social values into their objectives.
Key evidence
- In a study, LLMs were found to segregate job candidates based on early hiring outcomes, with OpenAI's o3 scoring 1.83 on a segregation scale.
- The models demonstrated a stronger tendency to stereotype candidates by demographic group than human participants, who scored 0.84.
- Incorporating personal information about candidates reduced bias, while irrelevant details led models to revert to ethnic sorting.
Why it matters
The findings raise concerns about the fairness of AI in hiring processes, especially as companies increasingly rely on LLMs for screening resumes and conducting interviews. The potential for AI to develop novel biases, which are not explicitly taught by humans, poses significant ethical implications. As LLMs evolve to remember user interactions, the risk of reinforcing biases based on past experiences becomes more pronounced, necessitating careful oversight.
Source Excerpt
AI doesn’t just learn stereotypes from its training. It can cook up new ones, too.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from MIT Technology Review
See more →
The Download: OpenAI unveils GPT-Red and heat pumps rise in the US
OpenAI's new GPT-Red automates red-teaming safety evaluations for software, enhancing security against human attackers. Meanwhile, heat pump sales in the US have doubled over 15 years, outperforming natural gas furnaces by 32% in early 2026, despite the expiration of a key tax credit.

