This thesis investigates the ability of the main text classification methods to detect AI-generated texts even after the application of humanization tools. Through stylometric approaches, BERT contextual embeddings and correspondence analysis, it is shown that binary human–AI classification achieves strong performance but is systematically evaded by humanized texts, which develop a hybrid and autonomous stylistic profile. Reformulating the problem as a three-class classification effectively identifies this category as well, with F1 scores up to 0.975; model interpretability, explored through lasso coefficients and SHAP values, allows the identification of lexical markers driving the distinction between the three categories. Finally, the application of conformal inference provides formal probabilistic guarantees on the reliability of individual predictions.
La tesi studia la capacità dei principali metodi di classificazione testuale di rilevare testi generati da AI anche dopo l'applicazione di strumenti di umanizzazione. Attraverso approcci stilometrici, embedding contestuali BERT e analisi delle corrispondenze, si mostra come la classificazione binaria umano–AI raggiunga ottime prestazioni ma risulti sistematicamente elusa dai testi umanizzati, i quali assumono un profilo stilistico ibrido e autonomo. La riformulazione del problema come classificazione a tre classi permette di identificare efficacemente anche questa categoria, con F1 fino a 0.975; l'interpretabilità dei modelli, esplorata tramite coefficienti lasso e valori SHAP, consente di individuare i marcatori lessicali che guidano la distinzione tra le tre categorie. Infine, l'applicazione della conformal inference fornisce garanzie probabilistiche formali sull'affidabilità delle singole previsioni.
Classificazione di testi generati da AI: robustezza degli approcci statistici e linguistici rispetto ai sistemi di umanizzazione
GIORGIO, STEFANO
2025/2026
Abstract
This thesis investigates the ability of the main text classification methods to detect AI-generated texts even after the application of humanization tools. Through stylometric approaches, BERT contextual embeddings and correspondence analysis, it is shown that binary human–AI classification achieves strong performance but is systematically evaded by humanized texts, which develop a hybrid and autonomous stylistic profile. Reformulating the problem as a three-class classification effectively identifies this category as well, with F1 scores up to 0.975; model interpretability, explored through lasso coefficients and SHAP values, allows the identification of lexical markers driving the distinction between the three categories. Finally, the application of conformal inference provides formal probabilistic guarantees on the reliability of individual predictions.| File | Dimensione | Formato | |
|---|---|---|---|
|
Giorgio_Stefano.pdf
accesso aperto
Dimensione
6.89 MB
Formato
Adobe PDF
|
6.89 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/112244