A Framework for Identifying, Categorizing, and Explaining Bias in AI-Generated Code
Source
Manaal Basha, Aimee M. Ribeiro, Gema Rodriguez-Perez
Author
Manaal Basha, Aimee M. Ribeiro, Gema Rodriguez-Perez
Date
Key takeaways · AI-distilled
The authors extended a dataset of biased AI-generated Python code with manual bias-category labels and human-written justifications, then used it as ground truth to test LLMs as bias detectors via in-context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → learning.
Gemini reached 80.14% classification accuracy with 84.0% precision and 95.7% recall; the best open model, Qwen3-coder, reached 82.45% accuracy but only 68.64% precision and 80.22% recall.
Model-written explanations scored about 80% similarity to the human justifications, and code-identification similarity ran 86% to 88%, which the authors read as substantial alignmentThe work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.Full definition → with expert reasoning.
Terms in this piece · Glossary
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
alignment — The work of making AI systems actually pursue what their builders and users intend, rather than something subtly or dangerously different.
Why it matters
A taxonomy-driven framework for identifying bias in AI-generated code finds Gemini reaches 80% classification accuracy against human-annotated ground truth, giving teams a concrete way to audit coding assistants for systematic bias.