Modern machine learning systems are increasingly deployed in high-stakes decision-making contexts, such as credit scoring, hiring, and insurance, where the people being evaluated are not passive observers. When individuals understand, or can estimate, how an algorithm works, they have rational incentives to strategically modify their data to obtain favourable outcomes. Standard classifiers, trained under the assumption that data is honest, are vulnerable to this behaviour. This thesis investigates the fragility of machine learning classifiers under strategic user behaviour, using game theory as the formal framework for modelling and simulating user actions. Starting from a synthetic mortgage dataset as a controlled baseline, three models of increasing complexity are developed: a static game of complete information in which a passive classifier faces a population of rational users who independently optimise their manipulation strategy; a static game of incomplete information introducing private user types and strategic deception, extended with a social learning mechanism through which users collectively refine their knowledge of the classifier over time; and a Stackelberg game in which the classifier, acting as the leader, anticipates the users’ rational response at design time. Each model generates a strategic dataset through Python simulation, which is then used to evaluate classifier performance across four experimental scenarios combining original and manipulated data. Results show that even simple strategic behaviour causes a 28.5% relative increase in false positive rate, that social learning amplifies this vulnerability by a further 24.8%, and that the Stackelberg classifier reduces vulnerability by 53.1% relative to the passive baseline.
Modern machine learning systems are increasingly deployed in high-stakes decision-making contexts, such as credit scoring, hiring, and insurance, where the people being evaluated are not passive observers. When individuals understand, or can estimate, how an algorithm works, they have rational incentives to strategically modify their data to obtain favourable outcomes. Standard classifiers, trained under the assumption that data is honest, are vulnerable to this behaviour. This thesis investigates the fragility of machine learning classifiers under strategic user behaviour, using game theory as the formal framework for modelling and simulating user actions. Starting from a synthetic mortgage dataset as a controlled baseline, three models of increasing complexity are developed: a static game of complete information in which a passive classifier faces a population of rational users who independently optimise their manipulation strategy; a static game of incomplete information introducing private user types and strategic deception, extended with a social learning mechanism through which users collectively refine their knowledge of the classifier over time; and a Stackelberg game in which the classifier, acting as the leader, anticipates the users’ rational response at design time. Each model generates a strategic dataset through Python simulation, which is then used to evaluate classifier performance across four experimental scenarios combining original and manipulated data. Results show that even simple strategic behaviour causes a 28.5% relative increase in false positive rate, that social learning amplifies this vulnerability by a further 24.8%, and that the Stackelberg classifier reduces vulnerability by 53.1% relative to the passive baseline.
Gaming the System: On the Fragility of Machine Learning under Strategic Behavior
SCARABELLO, LAURA
2025/2026
Abstract
Modern machine learning systems are increasingly deployed in high-stakes decision-making contexts, such as credit scoring, hiring, and insurance, where the people being evaluated are not passive observers. When individuals understand, or can estimate, how an algorithm works, they have rational incentives to strategically modify their data to obtain favourable outcomes. Standard classifiers, trained under the assumption that data is honest, are vulnerable to this behaviour. This thesis investigates the fragility of machine learning classifiers under strategic user behaviour, using game theory as the formal framework for modelling and simulating user actions. Starting from a synthetic mortgage dataset as a controlled baseline, three models of increasing complexity are developed: a static game of complete information in which a passive classifier faces a population of rational users who independently optimise their manipulation strategy; a static game of incomplete information introducing private user types and strategic deception, extended with a social learning mechanism through which users collectively refine their knowledge of the classifier over time; and a Stackelberg game in which the classifier, acting as the leader, anticipates the users’ rational response at design time. Each model generates a strategic dataset through Python simulation, which is then used to evaluate classifier performance across four experimental scenarios combining original and manipulated data. Results show that even simple strategic behaviour causes a 28.5% relative increase in false positive rate, that social learning amplifies this vulnerability by a further 24.8%, and that the Stackelberg classifier reduces vulnerability by 53.1% relative to the passive baseline.| File | Dimensione | Formato | |
|---|---|---|---|
|
Scarabello_Laura.pdf
accesso aperto
Dimensione
9.94 MB
Formato
Adobe PDF
|
9.94 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/110138