Large Language Models (LLMs) are increasingly being integrated into systems that influence decision-making processes. It is therefore crucial to investigate the extent to which they encode and reproduce harmful social stereotypes present in human-generated data. In this thesis, we conduct a multilingual, representation-level study of bias and refusal behaviour in the Llama-2-7B-Chat model and the GPT-OSS-20B model, using prompts in both English and Italian. We apply Contrastive Activation steering to build steering vectors for gender, race, religion, and refusal, using paired stereotype and anti-stereotype prompts. These prompts are derived from the StereoSet dataset and from custom bias-oriented prompts generated with GPT-4. A parallel Italian dataset is then created via machine translation and subsequently refined through human verification. We compare outputs for unsteered, bias-only-steered, and combined bias-and-refusal-steered responses across several model layers and steering strengths. The analysis investigates (i) whether stereotype and anti-stereotype prompts are linearly separable in activation space, (ii) whether suppressing refusal behaviour exposes biased continuations that safety mechanisms would otherwise conceal, and (iii) whether bias directions generalise across domains, models, and languages. The results provide evidence of social stereotypes in both models and languages, with particularly clear effects for gender and religion. They also indicate that bias can manifest differently across languages, even after Reinforcement Learning from Human Feedback (RLHF). Moreover, we find cases of language switching, in which models sometimes shift from Italian to English when responding to sensitive prompts. Taken together, these findings suggest that visible refusal should not be taken as proof that biased internal representations have disappeared. Instead, refusal may sometimes mask stereotypes encoded in the model, which can resurface under interventions such as activation steering. Notably, the study further indicates that RLHF tends to make the internal representations associated with different societal biases more similar, raising important questions about the extent to which the model distinguishes among different types of bias. Finally, this work offers practical insights into red-teaming LLMs via activation steering and, in particular, underscores the usefulness of incorporating an explicit refusal vector in such evaluations.

Large Language Models (LLMs) are increasingly being integrated into systems that influence decision-making processes. It is therefore crucial to investigate the extent to which they encode and reproduce harmful social stereotypes present in human-generated data. In this thesis, we conduct a multilingual, representation-level study of bias and refusal behaviour in the Llama-2-7B-Chat model and the GPT-OSS-20B model, using prompts in both English and Italian. We apply Contrastive Activation steering to build steering vectors for gender, race, religion, and refusal, using paired stereotype and anti-stereotype prompts. These prompts are derived from the StereoSet dataset and from custom bias-oriented prompts generated with GPT-4. A parallel Italian dataset is then created via machine translation and subsequently refined through human verification. We compare outputs for unsteered, bias-only-steered, and combined bias-and-refusal-steered responses across several model layers and steering strengths. The analysis investigates (i) whether stereotype and anti-stereotype prompts are linearly separable in activation space, (ii) whether suppressing refusal behaviour exposes biased continuations that safety mechanisms would otherwise conceal, and (iii) whether bias directions generalise across domains, models, and languages. The results provide evidence of social stereotypes in both models and languages, with particularly clear effects for gender and religion. They also indicate that bias can manifest differently across languages, even after Reinforcement Learning from Human Feedback (RLHF). Moreover, we find cases of language switching, in which models sometimes shift from Italian to English when responding to sensitive prompts. Taken together, these findings suggest that visible refusal should not be taken as proof that biased internal representations have disappeared. Instead, refusal may sometimes mask stereotypes encoded in the model, which can resurface under interventions such as activation steering. Notably, the study further indicates that RLHF tends to make the internal representations associated with different societal biases more similar, raising important questions about the extent to which the model distinguishes among different types of bias. Finally, this work offers practical insights into red-teaming LLMs via activation steering and, in particular, underscores the usefulness of incorporating an explicit refusal vector in such evaluations.

Activation Steering for Bias Detection in Large Language Models: A Study of Llama 2 and GPT-OSS-20B

CHAKRABORTHY, MEHULY
2025/2026

Abstract

Large Language Models (LLMs) are increasingly being integrated into systems that influence decision-making processes. It is therefore crucial to investigate the extent to which they encode and reproduce harmful social stereotypes present in human-generated data. In this thesis, we conduct a multilingual, representation-level study of bias and refusal behaviour in the Llama-2-7B-Chat model and the GPT-OSS-20B model, using prompts in both English and Italian. We apply Contrastive Activation steering to build steering vectors for gender, race, religion, and refusal, using paired stereotype and anti-stereotype prompts. These prompts are derived from the StereoSet dataset and from custom bias-oriented prompts generated with GPT-4. A parallel Italian dataset is then created via machine translation and subsequently refined through human verification. We compare outputs for unsteered, bias-only-steered, and combined bias-and-refusal-steered responses across several model layers and steering strengths. The analysis investigates (i) whether stereotype and anti-stereotype prompts are linearly separable in activation space, (ii) whether suppressing refusal behaviour exposes biased continuations that safety mechanisms would otherwise conceal, and (iii) whether bias directions generalise across domains, models, and languages. The results provide evidence of social stereotypes in both models and languages, with particularly clear effects for gender and religion. They also indicate that bias can manifest differently across languages, even after Reinforcement Learning from Human Feedback (RLHF). Moreover, we find cases of language switching, in which models sometimes shift from Italian to English when responding to sensitive prompts. Taken together, these findings suggest that visible refusal should not be taken as proof that biased internal representations have disappeared. Instead, refusal may sometimes mask stereotypes encoded in the model, which can resurface under interventions such as activation steering. Notably, the study further indicates that RLHF tends to make the internal representations associated with different societal biases more similar, raising important questions about the extent to which the model distinguishes among different types of bias. Finally, this work offers practical insights into red-teaming LLMs via activation steering and, in particular, underscores the usefulness of incorporating an explicit refusal vector in such evaluations.
2025
Activation Steering for Bias Detection in Large Language Models: A Study of Llama 2 and GPT-OSS-20B
Large Language Models (LLMs) are increasingly being integrated into systems that influence decision-making processes. It is therefore crucial to investigate the extent to which they encode and reproduce harmful social stereotypes present in human-generated data. In this thesis, we conduct a multilingual, representation-level study of bias and refusal behaviour in the Llama-2-7B-Chat model and the GPT-OSS-20B model, using prompts in both English and Italian. We apply Contrastive Activation steering to build steering vectors for gender, race, religion, and refusal, using paired stereotype and anti-stereotype prompts. These prompts are derived from the StereoSet dataset and from custom bias-oriented prompts generated with GPT-4. A parallel Italian dataset is then created via machine translation and subsequently refined through human verification. We compare outputs for unsteered, bias-only-steered, and combined bias-and-refusal-steered responses across several model layers and steering strengths. The analysis investigates (i) whether stereotype and anti-stereotype prompts are linearly separable in activation space, (ii) whether suppressing refusal behaviour exposes biased continuations that safety mechanisms would otherwise conceal, and (iii) whether bias directions generalise across domains, models, and languages. The results provide evidence of social stereotypes in both models and languages, with particularly clear effects for gender and religion. They also indicate that bias can manifest differently across languages, even after Reinforcement Learning from Human Feedback (RLHF). Moreover, we find cases of language switching, in which models sometimes shift from Italian to English when responding to sensitive prompts. Taken together, these findings suggest that visible refusal should not be taken as proof that biased internal representations have disappeared. Instead, refusal may sometimes mask stereotypes encoded in the model, which can resurface under interventions such as activation steering. Notably, the study further indicates that RLHF tends to make the internal representations associated with different societal biases more similar, raising important questions about the extent to which the model distinguishes among different types of bias. Finally, this work offers practical insights into red-teaming LLMs via activation steering and, in particular, underscores the usefulness of incorporating an explicit refusal vector in such evaluations.
LLM
Bias Evaluation
RLHF
Activation Steering
AI
File in questo prodotto:
File Dimensione Formato  
Chakraborthy_Mehuly.pdf

accesso aperto

Dimensione 3.62 MB
Formato Adobe PDF
3.62 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/110250