Narratives are a central object of study in the analysis of online news, yet the supervised pipelines that dominate the field depend on fixed, human-curated taxonomies of narrative labels. Building such taxonomies is expensive, and the resulting label sets are inevitably incomplete, which caps how far a purely supervised view can take us. This thesis asks a different question: how far can an unsupervised pipeline extract the narratives of a news corpus on its own, and can it surface coherent narratives that fall outside a given gold taxonomy? Working on the English portion of the PolyNarrative dataset, the dataset behind SemEval-2025 Task 10, we treat narrative extraction as a two-stage problem. A large language model first extracts short, paragraph-grounded claims under two complementary prompts, a direct claim prompt and a narrative-gap prompt motivated by the linguistic notion that narratives live in the inferences a reader is invited to supply across statements. The extracted claims are then embedded and clustered into a two-level structure that mirrors the coarse and fine levels of the gold taxonomy, without ever showing the model a gold label. We run a broad ablation over extraction prompts, four sentence embedders including a clustering-tuned Gemini model, and two clustering families, hierarchical agglomerative clustering and Leiden community detection, selecting the number of clusters with a leakage-free ensemble of internal validity criteria. Every configuration is evaluated against the gold taxonomy with an augmented-universe, multiplicity-aware BCubed suite together with Hungarian taxonomy alignment, both on the full corpus and under a stratified, article-level five-fold split. Error analysis shows that high-support narratives are recovered well, that the long Zipfian tail of rare narratives drives most of the residual error, and that several large and internally coherent clusters correspond to narratives that the gold taxonomy marks only as ”Other”, which is precisely the discovery behaviour we set out to test. We conclude that unsupervised narrative extraction is both a viable complement to taxonomy-bound classification and a practical tool for expanding narrative taxonomies.
UNSUPERVISED EXTRACTION OF NARRATIVES FROM NEWS ARTICLES
GUARINO, ANGELO
2025/2026
Abstract
Narratives are a central object of study in the analysis of online news, yet the supervised pipelines that dominate the field depend on fixed, human-curated taxonomies of narrative labels. Building such taxonomies is expensive, and the resulting label sets are inevitably incomplete, which caps how far a purely supervised view can take us. This thesis asks a different question: how far can an unsupervised pipeline extract the narratives of a news corpus on its own, and can it surface coherent narratives that fall outside a given gold taxonomy? Working on the English portion of the PolyNarrative dataset, the dataset behind SemEval-2025 Task 10, we treat narrative extraction as a two-stage problem. A large language model first extracts short, paragraph-grounded claims under two complementary prompts, a direct claim prompt and a narrative-gap prompt motivated by the linguistic notion that narratives live in the inferences a reader is invited to supply across statements. The extracted claims are then embedded and clustered into a two-level structure that mirrors the coarse and fine levels of the gold taxonomy, without ever showing the model a gold label. We run a broad ablation over extraction prompts, four sentence embedders including a clustering-tuned Gemini model, and two clustering families, hierarchical agglomerative clustering and Leiden community detection, selecting the number of clusters with a leakage-free ensemble of internal validity criteria. Every configuration is evaluated against the gold taxonomy with an augmented-universe, multiplicity-aware BCubed suite together with Hungarian taxonomy alignment, both on the full corpus and under a stratified, article-level five-fold split. Error analysis shows that high-support narratives are recovered well, that the long Zipfian tail of rare narratives drives most of the residual error, and that several large and internally coherent clusters correspond to narratives that the gold taxonomy marks only as ”Other”, which is precisely the discovery behaviour we set out to test. We conclude that unsupervised narrative extraction is both a viable complement to taxonomy-bound classification and a practical tool for expanding narrative taxonomies.| File | Dimensione | Formato | |
|---|---|---|---|
|
unsupervised-extraction-of-narratives-in-news-articles_Guarino-Angelo.pdf
Accesso riservato
Dimensione
1.69 MB
Formato
Adobe PDF
|
1.69 MB | Adobe PDF |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/110923