Wait, am I Being Fair?

Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG

Naihao Deng, Yilun Zhu, Joan Nwatu, Clayton Scott, Rada Mihalcea

University of Michigan · Language & Information Technologies (MichiganNLP)

⚠ This page and the underlying benchmarks contain examples of toxic and offensive stereotypes, used for the purpose of studying and mitigating bias.

Overview of deductive stereotyping and Fair-GCG reasoning-time steering
Deductive stereotyping: models apply population-level regularities to individual cases, producing logically coherent yet socially biased inferences. Fair-GCG discovers reasoning-time injection phrases that steer the model back toward fairness-aware reasoning.

Abstract

While reasoning generally improves fairness in recent large language models (LLMs), failures persist. In this work, we identify a failure mode, deductive stereotyping, in which models apply population-level statistical regularities to individual cases, producing logically coherent yet socially biased inferences. We provide a statistical interpretation of this phenomenon. To steer models toward fairness-aware reasoning, we propose a reasoning-time injection framework. We further introduce Fair-GCG to systematically discover effective injection phrases. Injection phrases discovered by Fair-GCG improve performance across multiple fairness benchmarks, generalize from smaller to larger LLMs, improve reasoning-level fairness, reduce bias in open-ended generation, and transfer to real-world fairness-sensitive tasks.

Key contributions

BibTeX

@article{deng2026fair,
  title   = {Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG},
  author  = {Deng, Naihao and Zhu, Yilun and Nwatu, Joan and Scott, Clayton and Mihalcea, Rada},
  journal = {arXiv preprint},
  year    = {2026}
}