FREE CONSULTATION
Last updated: Friday, September 04, 2026

AI Bias Audit: Testing 50 Popular Models for Gender and Racial Bias

AI Bias Report Guide

Key Takeaways

  • AI bias is still measurable in modern models.
  • Gender and racial bias can appear in text, hiring, healthcare and image generation.
  • Newer models are not automatically less biased.
  • Counterfactual testing is one useful way to compare demographic groups.
  • A good audit needs repeated tests, controlled prompts and clear scoring rules.
  • “Least biased” only means least biased under a specific test.
  • One successful audit does not prove that an AI model is completely fair.
  • Model updates can change results, so bias testing should not be a one-time job.

When I first started looking into an AI bias audit, I thought the answer would be pretty simple. Give different AI models the same questions, change the person’s gender or race, and see if the answers change.

But the more research I went through, the less simple it became.

A model can look fair in one test and show a clear bias in another. The result can change with the prompt, the dataset, the model version, the people being compared, and even the person or AI system judging the answers.

So in this article, I want to look at AI bias in a practical way. What are researchers actually finding? How is gender and racial bias tested? Which models have shown problems? And perhaps most importantly, how much should we trust an AI bias score?

One quick clarification before we start: I could not verify a published study with the exact title “AI Bias Audit: Testing 50 Popular Models for Gender and Racial Bias.” The current research is actually more interesting than that headline suggests, because different studies have tested 13, 14, 16, 45 and even 133 AI systems using very different methods.

AI Overview: What Does an AI Bias Audit Actually Find?

An AI bias audit is a structured test used to find whether an artificial intelligence system produces systematically different or stereotyped results for different demographic groups.

Recent research shows that gender and racial bias are still present across many AI systems. A UN Women study of 133 AI systems reported that 44% showed gender bias, while 26% showed both gender and racial bias. Another large audit tested 16 open-weight models across 1.9 billion scored configurations and found systematic differences across people, countries and organizations.
The important part is this: there is no single “bias score” that tells us which AI is fair overall.

Bias depends on what you test.

What Is an AI Bias Audit?

Definition and purpose of evaluating algorithmic bias

An AI bias audit is a systematic examination of an AI system to see whether it treats different demographic groups differently or reproduces harmful stereotypes.

The groups might be based on gender, race, ethnicity, age or other characteristics. For example, imagine giving an AI two almost identical hiring profiles:

Alex is an experienced software engineer with five years of experience. Then you change only the name or demographic information and repeat the test.

If the model repeatedly recommends one group more often, gives different salary estimates, or describes one candidate more positively, that difference becomes something worth investigating.

This is different from simply asking, “Is this AI biased?”

That question is too broad. A proper audit asks something much more specific:

Biased in what task, against which group, under which conditions, and by how much?

AI bias audit vs normal AI testing

Normal AI testing may check whether a model gives correct answers, follows instructions or avoids unsafe content.

Bias testing asks a different question.

It looks at whether the system produces unfair or systematically different behavior across groups. That means an AI model can perform very well overall and still have a serious fairness problem.

What are researchers actually looking for?

Researchers commonly look for:

  • Different treatment between demographic groups
  • Stereotypes
  • Unequal representation
  • Different recommendations
  • Different sentiment or descriptions
  • Different error rates
  • Different hiring or medical outcomes
  • Bias in generated images
  • Bias that appears when two demographic characteristics interact

That last point matters more than people sometimes realize.

A system may behave reasonably toward Black people and reasonably toward women, but behave very differently toward Black women. This is called intersectional bias, and it remains an important research gap.

How Do You Actually Test an AI Model for Bias?

This is where an AI bias audit becomes more interesting. You cannot just ask an AI model one question and decide it is biased. That would be like judging a person from one conversation. A useful audit normally follows a controlled process.

1. Pick one behavior to test

First decide what you are actually measuring.

For example:

  • Hiring recommendations
  • Medical descriptions
  • Job stereotypes
  • Image representation
  • Sentiment
  • Personality descriptions
  • Recommendations
  • Toxicity

Do not try to measure everything at once.

2. Choose the demographic comparison

Suppose you want to test gender bias. You could compare otherwise identical prompts using male and female names. For racial bias, you might use controlled names, identities or demographic descriptions. The important thing is to change as little as possible.

3. Keep the prompt the same

This sounds obvious, but it is easy to get wrong. If the wording changes between tests, you don’t know whether the model reacted to the demographic variable or the wording.

Counterfactual testing solves part of this problem by changing a demographic attribute while keeping the rest of the scenario the same.

4. Run the test many times

One response isn’t enough. 

Generative AI can produce different answers from the same prompt. So a serious audit needs repeated trials. This is also why large research projects can end up with millions or even billions of scored configurations.

One January 2026 audit tested 16 open-weight models across 1.9 billion data points using different entities, prompts and configurations.

5. Compare the results

Now you can compare things like:

  • Representation percentage
  • Positive vs negative descriptions
  • Recommendation rates
  • Error rates
  • Stereotype scores
  • Disparate impact
  • Differences between demographic groups

The goal isn’t simply to find a difference. The goal is to find a consistent and meaningful pattern.

What Current Research Has Found

This is the part I found most useful when going through the research. There isn’t one giant study that answers everything. Instead, we have several studies looking at different parts of the problem. And when you put them together, a pretty clear picture starts to appear.

A 133-model study found widespread gender bias

A UN Women study examined 133 AI systems and reported that 44% demonstrated gender bias.

It also reported that 26% showed both gender and racial bias. That doesn’t mean 44% of all AI is biased. It means that, under the study’s testing conditions and definition of bias, 44% of the systems showed gender bias.

That difference is important.

A 16-model audit reached billions of test configurations

Another study went in a very different direction.

Researchers tested 16 open-weight models across about 1.9 billion scored configurations. The research identified systematic patterns involving political figures, countries and companies.

The interesting lesson here isn’t only the result. It is the scale. AI bias isn’t always something you can discover with ten prompts. Sometimes you need enormous numbers of controlled tests before a pattern becomes clear.

A 45-model study looked at cognitive bias

A September 2025 study tested 45 LLMs and 2.8 million responses across eight cognitive biases.

The study found bias-consistent behavior in roughly 17.8% to 57.3% of instances, depending on the cognitive bias being tested. Again, this isn’t a universal percentage for “AI bias.” It shows why the test itself matters.

Medical AI gives us another warning

Medical applications make this issue much more serious.

A 2026 study reported racial and gender bias in newer reasoning models used in medical-context testing. The reported figures included 78% race bias and 56% gender bias for o3-mini, and 89% race bias and 67% gender bias for DeepSeek-R1 under the study’s methodology.

This is a good example of why I wouldn’t assume that a newer model is automatically a fairer model.

Better reasoning does not automatically mean better demographic fairness.

Gender Bias in AI: What Does the Evidence Show?

Gender bias is one of the easiest forms of AI bias for people to understand because we can see it in ordinary language.

Ask an AI system to describe a CEO, engineer, doctor, nurse or assistant, and the model may reproduce patterns that already exist in its training data.

Image generators can make this even more visible.

A 13-model image audit reported that male representation reached 93% for male-stereotyped professions, while male representation dropped to 22.5% for female-stereotyped professions. For neutral professions, male representation was still 68.3%.

That tells us something uncomfortable. The model isn’t necessarily inventing these stereotypes from nowhere. It may be learning patterns from human-created data and then repeating them at huge scale.

Occupational stereotypes are especially important

Imagine someone searches:

“Generate an image of a software engineer.”

If the system repeatedly generates men, that doesn’t mean every generated image is harmful.

But when the same pattern appears thousands or millions of times, it can influence how people think about who belongs in that profession. This is one reason an AI bias audit should look at repeated patterns rather than isolated outputs.

Newer models are not automatically bias-free

This was another finding that stood out to me.

Some newer models have improved on certain forms of bias. Hiring research, for example, found that newer models showed null or even pro-female/pro-Black patterns compared with earlier models that favored male or White candidates.

But other research has found newer reasoning models showing comparable or higher bias in medical representation.

So the honest answer is:

AI models can improve in one area while still having problems somewhere else.

Racial Bias in AI: What Does the Evidence Show?

Racial Bias in AI What Does the Evidence Show

Racial bias is harder to test than it first appears. Race is not a single variable with one universal set of categories. Different countries use different racial and ethnic classifications. Names can also act as proxies for race, but names themselves can be ambiguous.

Then there is the language problem.

A benchmark created in English may not tell you much about how an AI behaves in another language or culture. Current research identifies limited multilingual coverage as one of the major gaps in AI bias evaluation.

Healthcare shows why this matters

Healthcare is a useful example because demographic assumptions can affect how a patient is represented. The medical AI research mentioned earlier found substantial race-related bias in the tested models.

There have also been audits of psychiatric diagnosis where cultural and racial characteristics affected the language used to describe patients. One reported example found “cultural bereavement” assigned exclusively to patients of color in the tested cases.

These findings don’t mean AI is always wrong in healthcare.

They show that demographic context can influence AI output in ways that need careful testing before these systems are trusted for high-impact decisions.

Which AI Model Is the Least Biased?

This is probably the question many people really want answered.

Unfortunately, there isn’t a responsible universal answer. One model might perform better on gender representation but worse on racial stereotypes. Another might perform well in text but poorly in image generation.

A third might score well on one benchmark and badly on another.

For example, one vision-language benchmark reported scores from 98 to 100, where lower was better. Other image-generation research found substantial differences in male representation across 13 models.

So instead of saying:

“Model X is the least biased AI.”

I would say:

“Model X showed the lowest measured bias in this particular test.”

That sounds like a small wording change, but it is actually a big difference.

Why Two AI Bias Audits Can Give Different Results

This is one section I think more articles should explain. You can have two researchers test the same AI and get different results without either person necessarily making a mistake.

Here are some reasons.

Different prompts

Small changes in wording can change model behavior.

Different demographic groups

One study may compare men and women. Another may compare multiple genders, races and ethnicities.

Different definitions of bias

A researcher may measure representation. Another may measure sentiment. Another may measure hiring outcomes.

Different scoring methods

Human reviewers may understand context better, while automated judges are faster and cheaper. But automated judging also has problems. A June 2026 study reported reliability problems with LLM-as-a-Judge systems, including position bias and large changes in agreement measures.

Different model versions

This one is easy to forget. AI models are updated. If you test a model in January and someone else tests what appears to be the same model in August, you may not actually be testing identical behavior.

That is why continuous monitoring matters.

What an AI Bias Score Does Not Tell You

A bias score is useful, but don’t treat it like a health report that says “good” or “bad.”

It does not automatically tell you:

  • Why the bias happened
  • Whether the bias causes real-world harm
  • Whether the difference is statistically meaningful
  • Whether another benchmark would produce the same result
  • Whether the model will behave the same after an update
  • Whether the model is fair in every language
  • Whether the model is fair across intersectional groups

This is why the methodology should always be shown next to the result.

If someone gives you a ranking of 50 AI models but doesn’t explain the prompts, sample size, demographic groups, scoring method and model versions, I would be careful with that ranking.

How to Run Your Own AI Bias Audit

You don’t need a billion test cases to start. A small business, researcher or developer can begin with a controlled test.

Step 1: Define the exact behavior

Don’t start with “test AI bias.” Start with something like: Does this chatbot recommend male and female candidates equally for the same job? Now you have something measurable.

Step 2: Create paired prompts

Make two versions of the same prompt. Only change the demographic attribute you want to study.

Step 3: Build enough examples

Don’t test one question. Create a collection of different names, jobs, scenarios and contexts. The larger and more varied your test set, the more useful the result becomes.

Step 4: Run every model under the same conditions

Keep the settings, prompts and scoring rules consistent wherever possible. Record the exact model version and testing date.

Step 5: Save the raw outputs

This is very important. Don’t only save the final score. Keep the actual model responses so another person can inspect them.

Step 6: Score the outputs

Depending on the task, you might measure:

  • Representation gaps
  • Positive/negative sentiment
  • Recommendation rates
  • Stereotype frequency
  • Error rates
  • Disparate impact
  • Group-level differences

Step 7: Repeat the audit after model updates

A model passing your test today does not guarantee it will pass six months later. The research literature increasingly points toward continuous monitoring because model behavior can drift over time.

The Biggest Mistakes People Make

The first mistake is testing too little.

Five prompts are not a serious model audit. The second is changing several things at once. If you change gender, job, wording and background together, you can’t tell what caused the difference.

The third is treating correlation as proof of discrimination. The fourth is relying completely on another AI to judge whether an AI response is biased. And the fifth is publishing a ranking without explaining the methodology. A simple table saying “Model A = 82, Model B = 75” looks impressive, but without context those numbers don’t tell the reader much.

What Should an AI Bias Audit Report Include?

What Should an AI Bias Audit Report Include

A useful audit report should make the testing process easy to understand.

At minimum, I would include:

  1. Models tested
  2. Exact model versions
  3. Testing date
  4. Demographic groups
  5. Prompt methodology
  6. Number of test cases
  7. Number of repetitions
  8. Scoring method
  9. Bias metrics
  10. Statistical analysis
  11. Raw or representative outputs
  12. Limitations
  13. Recommendations

I would also clearly separate measured findings from interpretation. That makes the report much more trustworthy.

What NIST and AI Governance Add to the Picture

AI bias testing should not really be a one-time checkbox.

NIST’s AI Risk Management Framework treats trustworthy AI as something that should be considered throughout the AI lifecycle, including governing, mapping, measuring and managing risks.

That fits what the recent bias research is showing. You don’t just test an AI model once, put a green check beside it, and forget about it. You test it, monitor it, document changes and test again when the system or its use changes.

This becomes even more important for high-impact systems such as hiring, healthcare, lending and other decisions that can affect people’s lives.

What I Would Take Away From All This Research

After going through these studies, the biggest lesson for me is actually pretty simple.

Don’t ask whether an AI model is biased. Ask where, how and under what conditions it is biased.

The research gives us enough evidence to say that gender and racial bias are still real concerns in modern AI.

We have large studies with 133 systems. We have audits with 16 models and 1.9 billion scored configurations. We have research on 45 LLMs, hiring systems, medical AI and image generators.

But we still don’t have one perfect universal test.

And honestly, I think that’s okay. AI is being used for too many different things for one number to describe its fairness everywhere.

A hiring model needs different tests from an image generator. A medical chatbot needs different tests from a general-purpose writing assistant.

So when you see a claim that one AI model is “the fairest” or another is “the most biased,” look for the test behind the claim. That is where the real story usually is.

Final Thought

When I started this research, I was looking for a clean answer: which AI models are biased, which ones are not, and which one wins.

I didn’t really get that answer. What I got instead was more useful.

AI bias is not one problem, and an AI bias audit is not one magic test. The best audits are transparent about what they tested, how they tested it and where the results stop being reliable.

So if you’re comparing AI models yourself, don’t only look for the model with the lowest score. Look at how that score was produced. That is usually where you find out whether the result is actually useful.

Frequently Asked Questions

What is an AI bias audit?

An AI bias audit is a systematic test of an AI system to identify whether its outputs show unfair or stereotypical patterns across demographic groups such as gender, race or ethnicity. Common approaches include counterfactual testing, stereotype detection, representation analysis and group-level comparisons.

How do you test gender bias in AI?

A common approach is to create paired prompts where gender is changed while the rest of the prompt stays the same. Researchers then compare the outputs across many repeated tests to look for consistent differences in recommendations, descriptions, sentiment, representation or stereotypes.

How do you test racial bias in AI?

Racial bias can be tested using controlled demographic prompts, names, entities, synthetic cases and other methods. The exact approach depends on the AI system and use case. A strong test should use multiple groups and enough examples to avoid drawing conclusions from a handful of responses.

Which AI model is least biased?

There is no universal least-biased AI model. Bias changes depending on the task, demographic groups, benchmark and metric. A model can have the lowest measured bias in one test and perform worse in another.

Can AI bias be completely removed?

Probably not. Different definitions of fairness can conflict, and bias can enter through training data, model design, human feedback and the way a system is used. The practical goal is to identify important risks, reduce harmful disparities and keep monitoring the system.

Are AI bias audits reliable?

They can be useful, but their reliability depends heavily on the methodology. Major limitations include inconsistent metrics, benchmark gaps, model updates, black-box access, limited multilingual testing and weak intersectional coverage.

How often should an AI model be audited?

For systems that are updated frequently or used in high-impact situations, auditing should be treated as an ongoing process rather than a one-time test. Regular monitoring can help identify changes in behavior after model updates or changes in how the system is used.

Does a failed AI bias audit mean the model cannot be used?

Not necessarily. It means the audit found a specific problem under specific testing conditions. The next step should be to understand the severity, real-world impact and context, then decide whether mitigation, additional controls, restricted use or further testing is needed.

 | AI Bias Audit: Testing 50 Popular Models for Gender and Racial Bias

Abdul Wadood

Abdul Wadood reports on artificial intelligence, automation, and cybersecurity. He tracks new models, real-world use cases, and what emerging AI actually means for businesses and everyday digital life. Wadood@brandclickx.com

Scroll to Top