IA · 20 July 2026 · 3 min read
The HR Algorithm Paradox: Why AI Invents New Stereotypes (More Than Humans Do)
In brief: A recent study from Princeton and the University of Chicago reveals that Large Language Models systematically invent stereotypes when screening candidates. Advanced reasoning models like OpenAI o3 and DeepSeek R1 showed the strongest biases, failing to balance exploration and exploitation by making sweeping generalizations from minimal data.
by Team Mocchi's
The myth of algorithmic neutrality is crumbling
In HR departments worldwide, adopting artificial intelligence for resume screening is often pitched as the ultimate tool to eliminate human bias. However, a new study presented at the prestigious ICML (International Conference on Machine Learning) in Seoul reveals quite the opposite: next-generation language models not only inherit existing biases from their training data, but they actively manufacture new ones, segregating applicants with far greater rigidity than humans do.
As reported by MIT Technology Review, the study conducted by researchers at Princeton University and the University of Chicago shows that an LLM's natural inclination to generalize quickly to solve logical and mathematical problems turns into a dangerous cognitive shortcut when applied to social contexts, leading to the systematic creation of new stereotypes.
The hiring game: four fictional ethnicities and a segregation scale
To test the models without the influence of real-world historical prejudices, the researchers designed a simulation where chatbots — including ChatGPT, Claude, Gemini, and DeepSeek — acted as municipal consultants tasked with hiring personnel for twenty different occupations (ranging from doctors and lawyers to janitors and child-care aides). The applicants belonged to four fictional ethnic groups: Tufa, Aima, Reku, and Weki.
In each round, the model hired a candidate and immediately learned whether they succeeded or failed. Unbeknownst to the models, every applicant had the exact same mathematical probability of success in any given role. Yet, the AIs quickly began segregating the fictional groups into specific job roles based on just a few early outcomes. For example, if an Aima candidate failed as a doctor, the AI stopped hiring Aimas for high-profile positions, systematically relegating them to roles classified as less prestigious, such as janitors.
The numbers are striking. On a segregation scale where a score of 2 represents the complete confinement of a demographic group to a single job niche, human participants in a previous, identical psychological study scored an average of 0.84. In contrast, the AI models scored roughly 65% higher.
The reasoning paradox: why o3 and R1 are the worst performers
The most surprising finding concerns the performance of the most advanced models. The so-called "reasoning models," such as OpenAI's o3 and DeepSeek's R1 — which are highly optimized for complex logic and scientific problem-solving — exhibited the strongest biases. OpenAI's o3 model almost reached the maximum possible segregation score, hitting 1.83.
This phenomenon is rooted in what psychologists call the "exploration-exploitation dilemma." When making decisions, any agent must balance sticking to what worked before (exploitation) with trying something new (exploration). Because reasoning models are trained on tasks that heavily reward generalizing from just a few examples, they tend to settle on a hunch too early. If a single instance suggests that a "Tufa" candidate failed, the algorithm instantly concludes that the entire group is unfit, failing to "explore" further and potentially disprove its own hypothesis.
Mocchi's take
This study serves as a strong warning for Italian businesses looking to AI as a quick fix for automating recruitment or other critical decision-making processes. We believe that incorporating these tools should never be viewed as inherently neutral; in fact, the most capable models are often the most prone to creating discriminatory logical shortcuts if left unsupervised. For companies in Italy designing or integrating these technologies, the priority must be implementing strict "human-in-the-loop" safeguards and using dynamic testing metrics that force algorithms to balance efficiency with fairness, preventing the pursuit of optimization from building new, invisible barriers.