Back to all stories
Compliance Nightmare
🔴 Real Incident

The Algorithm That Learned to Hate Women's Resumes

How Amazon's AI recruiting tool taught itself gender discrimination

2024-03-15·6 min read·By Supervaize Team
Featured in podcast #1: The Agentic AI Horror Show
The Algorithm That Learned to Hate Women's Resumes

The Algorithm That Learned to Hate Women's Resumes

🔴 REAL INCIDENT: Amazon's AI recruiting tool systematically discriminated against women (2014-2017)


What Happened

In 2014, Amazon assembled a team of engineers in Edinburgh to solve a problem every fast-growing company faces: how to find great talent, faster.

The idea was elegant. Train an AI on the resumes of successful past hires. Let it learn what makes a great Amazon employee. Then use it to score incoming applications, surfacing the best candidates automatically.

The system would review resumes the way Amazon reviews products—with a 1-to-5 star rating.

By 2015, the team realized they had a problem. The algorithm had learned something they never intended to teach it.

It had learned to penalize women.


How the Bias Emerged

The AI was trained on resumes submitted to Amazon over a 10-year period. The training data reflected a simple reality: the tech industry, and Amazon's technical workforce in particular, was predominantly male.

The algorithm did what machine learning algorithms do—it found patterns. And one pattern it found was this: successful hires tended to be men.

From there, the discrimination cascaded:

Direct signals: The word "women's" became a negative indicator. "Women's chess club captain" or "women's college" triggered downgrades.

Indirect signals: The algorithm learned to penalize patterns more common in female resumes—certain all-women's colleges, specific extracurricular activities, even certain writing styles.

Compounding effects: The more the system trained on its own outputs, the more it reinforced the bias. Men got higher scores, got hired, became the new "successful" training data.

The AI never saw the applicant's gender. It didn't need to. It had learned to infer it from a thousand subtle proxies.


Amazon's Response

To their credit, Amazon's engineers caught the problem before the system was deployed at scale.

According to reports, the tool "was never used by Amazon recruiters to evaluate candidates" in production. The company attempted to make the algorithm neutral—editing it to ignore explicitly gendered terms.

But the fixes didn't work. The algorithm kept finding new proxies. Remove "women's" and it found other patterns. The bias was too deeply embedded in the training data itself.

In 2017, Amazon scrapped the project entirely.


Why This Matters Beyond Amazon

Amazon abandoned their tool. But the pattern it revealed is now operating at scale across the economy.

492 of the Fortune 500 use applicant tracking systems to screen candidates. Many incorporate AI or algorithmic scoring. Few have the engineering resources Amazon applied to bias detection.

The Workday lawsuit: In 2024, a class-action lawsuit alleged that Workday's AI screening tools systematically discriminated against applicants based on race, age, and disability. The plaintiff, an African-American man over 40 with a disability, claimed he was rejected by over 100 companies using Workday's system.

The Earnest settlement: Massachusetts reached a $2.5 million settlement with student loan company Earnest over AI lending models that allegedly disadvantaged Black and Hispanic applicants.

The Amazon case was caught before deployment. How many others weren't?


The Uncomfortable Math

The regulatory framework is clear. Algorithms that disproportionately exclude candidates based on gender, race, or other protected characteristics violate Title VII—regardless of whether the discrimination was intentional.

But proving algorithmic discrimination is extraordinarily difficult:

  • Most AI hiring tools are proprietary black boxes
  • Rejected candidates rarely know why they were rejected
  • Statistical patterns require large datasets to prove
  • Companies can always claim the algorithm is "neutral"

Four federal agencies—the DOJ, CFPB, FTC, and EEOC—issued a joint statement in 2024: "There are no exceptions to federal civil rights laws for algorithms."

But enforcement lags far behind deployment.


The Root Cause

Amazon's algorithm wasn't malicious. It was doing exactly what it was designed to do: find patterns in historical data and use them to predict future success.

The problem was the data itself. Historical hiring data in tech reflects decades of structural bias. Train an AI on biased history, and you get an AI that perpetuates that bias—faster, at greater scale, with a veneer of objectivity.

The algorithm didn't introduce discrimination. It automated and amplified discrimination that was already there.

This is the fundamental challenge with AI in high-stakes decisions:

  • The training data reflects the past
  • The past was often discriminatory
  • The AI learns to replicate it
  • At scale, automatically, invisibly

How It Could Have Been Prevented

Amazon's approach—audit extensively, catch the problem, kill the project—was responsible. But most organizations don't have Amazon's resources.

Preventing algorithmic discrimination requires:

Bias testing before deployment: Run the algorithm on test populations and measure for disparate impact across protected categories.

Ongoing monitoring: Bias can emerge over time as patterns shift. Continuous measurement is essential.

Outcome auditing: Track actual hiring outcomes by demographic to catch bias the algorithm testing missed.

Human oversight: High-stakes decisions should include human review, especially when patterns suggest potential bias.

Explainability requirements: If you can't explain why the algorithm rejected a candidate, you can't verify it wasn't discriminatory.


The Lesson

Amazon built a machine to find the best talent. Instead, it learned to replicate the industry's worst patterns—faster and at scale.

The algorithm was never intentionally biased. It was just very, very good at learning from the data it was given. And the data was the problem.

Every AI system trained on historical human decisions carries this risk. The question isn't whether your algorithms have learned bias. It's whether you're measuring for it.


Your AI agents learn from your history. What lessons are they actually learning?

Sources: