Symbolic Computation Group

David R. Cheriton School of Computer Science
University of Waterloo, Waterloo, Ontario, Canada

Tuesday, August 25, 2026, at the University of Waterloo
Inferred Author Gender as a Variable Affecting LLM Behaviour
Avery Hiebert, Master's candidate, University of Waterloo
Supervisor: Professor George Labahn

Abstract:

The presence of undesirable bias in Large Language Models (LLMs) and other NLP systems has been studied extensively. However, the widespread framing of “bias” as a distinct phenomenon separable from “legitimate” uses of gender has been identified as a weakness of the existing literature, as has a lack of engagement with the meaning and use of “gender” as a category.

We investigate the modeling of “author gender” as a source of gender-influenced behaviour in GPT-2, using an exploratory approach based on interpretability techniques rather than the traditional “biased/unbiased” dichotomy. A simple zero-shot prompt reveals that GPT-2 has learned to estimate author gender in some contexts. Attribution experiments reveal stereotypical associations, both stylistic and semantic, that inform the model’s estimation of author gender. Causal intervention experiments using a linear probing classifier suggest that author gender information influences model output, including reflecting stereotypical gender associations. These results have implications for the detection and mitigation of gender bias and the broader understanding of how gender is “used” by LLMs.

 

Last modified on Tuesday, 18 August 2026, at 22:47 hours.