Topic modeling is a family of unsupervised learning techniques that allows for the automatic identification of latent themes within a textual corpus. In this work, such techniques are applied to the digital communication of eight Italian science museums on social media platforms, with the aim of extracting and comparing the recurring themes present in the published content. The classical models Latent Dirichlet Allocation (LDA) and Structural Topic Model (STM) are first implemented, serving as a baseline for subsequent analyses. BERTopic is then introduced, a modern approach to topic modeling that, unlike classical bag-of-words methods, leverages contextual document representations obtained through transformer models. The results are subsequently compared with those of LDA and STM in terms of coherence, exclusivity and diversity. Several variants of the BERTopic pipeline are further explored, with the aim of assessing possible improvements on the Italian corpus under examination. The modifications concern the dimensionality reduction algorithm, the clustering method and the text representation strategy. Finally, an integration of Gibbs sampling within the BERTopic pipeline is tested, with the intent of refining document-topic assignments by inheriting the inferential robustness of classical probabilistic models.

Il topic modeling è una famiglia di tecniche di apprendimento non supervisionato che permette di identificare automaticamente i temi latenti presenti in un corpus testuale. In questo lavoro, tali tecniche vengono applicate alla comunicazione digitale di otto musei scientifici italiani su piattaforme social, con l'obiettivo di estrarre e confrontare i macro-temi presenti nei contenuti pubblicati. Vengono inizialmente implementati i modelli classici Latent Dirichlet Allocation (LDA) e Structural Topic Model (STM), che costituiscono una baseline di riferimento per le analisi successive. Successivamente viene introdotto BERTopic, un approccio moderno al topic modeling che, a differenza dei metodi classici basati su bag-of-words, sfrutta rappresentazioni contestuali dei documenti ottenute tramite modelli transformer. I risultati vengono quindi confrontati con quelli di LDA e STM in termini di coerenza, esclusività e diversità. In seguito vengono esplorate diverse varianti della pipeline di BERTopic, con l'obiettivo di valutarne possibili miglioramenti sul corpus italiano in esame. Le modifiche riguardano l'algoritmo di riduzione dimensionale, il metodo di clustering e la strategia di rappresentazione testuale. Viene infine testata un'integrazione del Gibbs sampling all'interno della pipeline di BERTopic, con l'intento di raffinare le assegnazioni documento-topic ereditando la solidità inferenziale dei modelli probabilistici classici.

Topic Modeling applicato alla comunicazione digitale di musei scientifici italiani: implementazione, limiti e miglioramenti di BERTopic per la lingua italiana

SIMONETTO, ANDREA
2025/2026

Abstract

Topic modeling is a family of unsupervised learning techniques that allows for the automatic identification of latent themes within a textual corpus. In this work, such techniques are applied to the digital communication of eight Italian science museums on social media platforms, with the aim of extracting and comparing the recurring themes present in the published content. The classical models Latent Dirichlet Allocation (LDA) and Structural Topic Model (STM) are first implemented, serving as a baseline for subsequent analyses. BERTopic is then introduced, a modern approach to topic modeling that, unlike classical bag-of-words methods, leverages contextual document representations obtained through transformer models. The results are subsequently compared with those of LDA and STM in terms of coherence, exclusivity and diversity. Several variants of the BERTopic pipeline are further explored, with the aim of assessing possible improvements on the Italian corpus under examination. The modifications concern the dimensionality reduction algorithm, the clustering method and the text representation strategy. Finally, an integration of Gibbs sampling within the BERTopic pipeline is tested, with the intent of refining document-topic assignments by inheriting the inferential robustness of classical probabilistic models.
2025
Topic modeling applied to the digital communication of Italian science museums: implementation, limitations, and improvements of BERTopic for the Italian language
Il topic modeling è una famiglia di tecniche di apprendimento non supervisionato che permette di identificare automaticamente i temi latenti presenti in un corpus testuale. In questo lavoro, tali tecniche vengono applicate alla comunicazione digitale di otto musei scientifici italiani su piattaforme social, con l'obiettivo di estrarre e confrontare i macro-temi presenti nei contenuti pubblicati. Vengono inizialmente implementati i modelli classici Latent Dirichlet Allocation (LDA) e Structural Topic Model (STM), che costituiscono una baseline di riferimento per le analisi successive. Successivamente viene introdotto BERTopic, un approccio moderno al topic modeling che, a differenza dei metodi classici basati su bag-of-words, sfrutta rappresentazioni contestuali dei documenti ottenute tramite modelli transformer. I risultati vengono quindi confrontati con quelli di LDA e STM in termini di coerenza, esclusività e diversità. In seguito vengono esplorate diverse varianti della pipeline di BERTopic, con l'obiettivo di valutarne possibili miglioramenti sul corpus italiano in esame. Le modifiche riguardano l'algoritmo di riduzione dimensionale, il metodo di clustering e la strategia di rappresentazione testuale. Viene infine testata un'integrazione del Gibbs sampling all'interno della pipeline di BERTopic, con l'intento di raffinare le assegnazioni documento-topic ereditando la solidità inferenziale dei modelli probabilistici classici.
Topic Modeling
BERTopic
LDA
STM
File in questo prodotto:
File Dimensione Formato  
Simonetto_Andrea.pdf

accesso aperto

Dimensione 782.07 kB
Formato Adobe PDF
782.07 kB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/112250