Can large language models generate symbolic music? This thesis explores this question using LilyPond notation as a testbed. Starting from 11,995 LilyPond files of Baroque instrumental music, we filter to 391 scores, split them into 1,888 movements, and extract 17,040 singlevoice assignments. A multi-stage normalization pipeline resolves includes, cleans syntax, and removes engraving-only constructs. We compare three approaches: zero-shot prompting, few-shot in-context learning, and LowRank Adaptation (LoRA) fine-tuning of the GPT-OSS-20B model. The training representation preserves LilyPond note and rest strings while augmenting them with domain-specific tokens for key, time, tempo, and voice. Evaluation combines compilation success, structural validity, musical metrics, and qualitative review. Zero-shot generation works for simple patterns but tends to produce short, mechanically regular material and degrades as musical structure becomes more complex. Few-shot prompting improves surface form, but its overall impact remains limited. Fine-tuning yields the most reliable behavior, producing substantially longer continuations that often compile and render successfully and more frequently exhibit musically coherent phrasing. This thesis contributes a preprocessing and normalization pipeline, a filtered and standardized dataset, and an evaluation framework for symbolic music generation.

I modelli linguistici di grandi dimensioni possono generare musica simbolica? Questa tesi esplora tale domanda utilizzando la notazione LilyPond come banco di prova. A partire da 11.995 file LilyPond di musica strumentale barocca, filtriamo 391 partiture, le suddividiamo in 1.888 movimenti ed estraiamo 17.040 assegnazioni a voce singola. Una pipeline di normalizzazione multi-stadio risolve gli include, pulisce la sintassi e rimuove i costrutti destinati esclusivamente all’incisione grafica. Confrontiamo tre approcci: prompting zero-shot, apprendimento in-context few-shot e finetuning tramite Low-Rank Adaptation (LoRA) del modello GPT-OSS-20B. La rappresentazione di addestramento preserva le stringhe di note e pause di LilyPond, arricchendole con token specifici di dominio per tonalità, metro, tempo e voce. La valutazione combina il successo di compilazione, la validità strutturale, metriche musicali e revisione qualitativa. La generazione zero-shot funziona per schemi semplici, ma tende a produrre materiale breve e meccanicamente regolare e peggiora all’aumentare della complessità strutturale. Il prompting few-shot migliora la forma superficiale, ma il suo impatto complessivo resta limitato. Il finetuning fornisce il comportamento più affidabile, producendo continuazioni sostanzialmente più lunghe che spesso compilano e vengono renderizzate correttamente e che mostrano più frequentemente una fraseologia musicalmente coerente. Questa tesi contribuisce con una pipeline di preprocessing e normalizzazione, un dataset filtrato e standardizzato e un framework di valutazione per la generazione di musica simbolica.

Evaluating and Fine-Tuning Large Language Models for Symbolic Music Generation via Normalized LilyPond Notation

TORABI, MOHAMMAD
2025/2026

Abstract

Can large language models generate symbolic music? This thesis explores this question using LilyPond notation as a testbed. Starting from 11,995 LilyPond files of Baroque instrumental music, we filter to 391 scores, split them into 1,888 movements, and extract 17,040 singlevoice assignments. A multi-stage normalization pipeline resolves includes, cleans syntax, and removes engraving-only constructs. We compare three approaches: zero-shot prompting, few-shot in-context learning, and LowRank Adaptation (LoRA) fine-tuning of the GPT-OSS-20B model. The training representation preserves LilyPond note and rest strings while augmenting them with domain-specific tokens for key, time, tempo, and voice. Evaluation combines compilation success, structural validity, musical metrics, and qualitative review. Zero-shot generation works for simple patterns but tends to produce short, mechanically regular material and degrades as musical structure becomes more complex. Few-shot prompting improves surface form, but its overall impact remains limited. Fine-tuning yields the most reliable behavior, producing substantially longer continuations that often compile and render successfully and more frequently exhibit musically coherent phrasing. This thesis contributes a preprocessing and normalization pipeline, a filtered and standardized dataset, and an evaluation framework for symbolic music generation.
2025
Evaluating and Fine-Tuning Large Language Models for Symbolic Music Generation via Normalized LilyPond Notation
I modelli linguistici di grandi dimensioni possono generare musica simbolica? Questa tesi esplora tale domanda utilizzando la notazione LilyPond come banco di prova. A partire da 11.995 file LilyPond di musica strumentale barocca, filtriamo 391 partiture, le suddividiamo in 1.888 movimenti ed estraiamo 17.040 assegnazioni a voce singola. Una pipeline di normalizzazione multi-stadio risolve gli include, pulisce la sintassi e rimuove i costrutti destinati esclusivamente all’incisione grafica. Confrontiamo tre approcci: prompting zero-shot, apprendimento in-context few-shot e finetuning tramite Low-Rank Adaptation (LoRA) del modello GPT-OSS-20B. La rappresentazione di addestramento preserva le stringhe di note e pause di LilyPond, arricchendole con token specifici di dominio per tonalità, metro, tempo e voce. La valutazione combina il successo di compilazione, la validità strutturale, metriche musicali e revisione qualitativa. La generazione zero-shot funziona per schemi semplici, ma tende a produrre materiale breve e meccanicamente regolare e peggiora all’aumentare della complessità strutturale. Il prompting few-shot migliora la forma superficiale, ma il suo impatto complessivo resta limitato. Il finetuning fornisce il comportamento più affidabile, producendo continuazioni sostanzialmente più lunghe che spesso compilano e vengono renderizzate correttamente e che mostrano più frequentemente una fraseologia musicalmente coerente. Questa tesi contribuisce con una pipeline di preprocessing e normalizzazione, un dataset filtrato e standardizzato e un framework di valutazione per la generazione di musica simbolica.
Symbolic Music
Language Models
LilyPond
Normalization
Fine-Tuning
File in questo prodotto:
File Dimensione Formato  
Torabi_Mohammad.pdf

embargo fino al 14/04/2027

Dimensione 1.82 MB
Formato Adobe PDF
1.82 MB Adobe PDF

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/106863