Can large language models generate symbolic music? This thesis explores this question using LilyPond notation as a testbed. Starting from 11,995 LilyPond files of Baroque instrumental music, we filter to 391 scores, split them into 1,888 movements, and extract 17,040 singlevoice assignments. A multi-stage normalization pipeline resolves includes, cleans syntax, and removes engraving-only constructs. We compare three approaches: zero-shot prompting, few-shot in-context learning, and LowRank Adaptation (LoRA) fine-tuning of the GPT-OSS-20B model. The training representation preserves LilyPond note and rest strings while augmenting them with domain-specific tokens for key, time, tempo, and voice. Evaluation combines compilation success, structural validity, musical metrics, and qualitative review. Zero-shot generation works for simple patterns but tends to produce short, mechanically regular material and degrades as musical structure becomes more complex. Few-shot prompting improves surface form, but its overall impact remains limited. Fine-tuning yields the most reliable behavior, producing substantially longer continuations that often compile and render successfully and more frequently exhibit musically coherent phrasing. This thesis contributes a preprocessing and normalization pipeline, a filtered and standardized dataset, and an evaluation framework for symbolic music generation.
I modelli linguistici di grandi dimensioni possono generare musica simbolica? Questa tesi esplora tale domanda utilizzando la notazione LilyPond come banco di prova. A partire da 11.995 file LilyPond di musica strumentale barocca, filtriamo 391 partiture, le suddividiamo in 1.888 movimenti ed estraiamo 17.040 assegnazioni a voce singola. Una pipeline di normalizzazione multi-stadio risolve gli include, pulisce la sintassi e rimuove i costrutti destinati esclusivamente all’incisione grafica. Confrontiamo tre approcci: prompting zero-shot, apprendimento in-context few-shot e finetuning tramite Low-Rank Adaptation (LoRA) del modello GPT-OSS-20B. La rappresentazione di addestramento preserva le stringhe di note e pause di LilyPond, arricchendole con token specifici di dominio per tonalità, metro, tempo e voce. La valutazione combina il successo di compilazione, la validità strutturale, metriche musicali e revisione qualitativa. La generazione zero-shot funziona per schemi semplici, ma tende a produrre materiale breve e meccanicamente regolare e peggiora all’aumentare della complessità strutturale. Il prompting few-shot migliora la forma superficiale, ma il suo impatto complessivo resta limitato. Il finetuning fornisce il comportamento più affidabile, producendo continuazioni sostanzialmente più lunghe che spesso compilano e vengono renderizzate correttamente e che mostrano più frequentemente una fraseologia musicalmente coerente. Questa tesi contribuisce con una pipeline di preprocessing e normalizzazione, un dataset filtrato e standardizzato e un framework di valutazione per la generazione di musica simbolica.
Evaluating and Fine-Tuning Large Language Models for Symbolic Music Generation via Normalized LilyPond Notation
TORABI, MOHAMMAD
2025/2026
Abstract
Can large language models generate symbolic music? This thesis explores this question using LilyPond notation as a testbed. Starting from 11,995 LilyPond files of Baroque instrumental music, we filter to 391 scores, split them into 1,888 movements, and extract 17,040 singlevoice assignments. A multi-stage normalization pipeline resolves includes, cleans syntax, and removes engraving-only constructs. We compare three approaches: zero-shot prompting, few-shot in-context learning, and LowRank Adaptation (LoRA) fine-tuning of the GPT-OSS-20B model. The training representation preserves LilyPond note and rest strings while augmenting them with domain-specific tokens for key, time, tempo, and voice. Evaluation combines compilation success, structural validity, musical metrics, and qualitative review. Zero-shot generation works for simple patterns but tends to produce short, mechanically regular material and degrades as musical structure becomes more complex. Few-shot prompting improves surface form, but its overall impact remains limited. Fine-tuning yields the most reliable behavior, producing substantially longer continuations that often compile and render successfully and more frequently exhibit musically coherent phrasing. This thesis contributes a preprocessing and normalization pipeline, a filtered and standardized dataset, and an evaluation framework for symbolic music generation.| File | Dimensione | Formato | |
|---|---|---|---|
|
Torabi_Mohammad.pdf
embargo fino al 14/04/2027
Dimensione
1.82 MB
Formato
Adobe PDF
|
1.82 MB | Adobe PDF |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/106863