This thesis presents the design and implementation of a speech-to-MIDI framework for expressive automatic piano playback on a Yamaha Disklavier. The work investigates how human voice signals can be transformed into symbolic musical events through Digital Signal Processing (DSP), then reproduced acoustically through electromechanical piano actuation. The proposed system combines Short-Time Fourier Transform (STFT)-based time-frequency analysis, spectral peak detection, frequency-to-MIDI mapping, and velocity estimation to generate musically interpretable MIDI note streams from speech recordings. In addition to the core conversion pipeline, the project implements a modular software architecture with a FastAPI-based web workflow, including file upload, parameter validation, asynchronous job processing, status tracking, and MIDI download. This integration addresses both algorithmic and engineering requirements for practical use. The study is grounded in historical and artistic context, from player pianos to modern Disklavier systems, and is conceptually connected to speech-like mechanical piano practices such as Peter Ablinger’s Deus Cantando. Results demonstrate the feasibility of transforming monophonic speech into expressive symbolic piano performance, while highlighting current limitations related to transcription precision, non-stationary speech complexity, and future real-time.

Development of a Web-Based Software Platform for Converting Speech Input into Symbolic Music Notation Playable on a Disklavier Automatic Piano

FATAHI, GHAZAL
2025/2026

Abstract

This thesis presents the design and implementation of a speech-to-MIDI framework for expressive automatic piano playback on a Yamaha Disklavier. The work investigates how human voice signals can be transformed into symbolic musical events through Digital Signal Processing (DSP), then reproduced acoustically through electromechanical piano actuation. The proposed system combines Short-Time Fourier Transform (STFT)-based time-frequency analysis, spectral peak detection, frequency-to-MIDI mapping, and velocity estimation to generate musically interpretable MIDI note streams from speech recordings. In addition to the core conversion pipeline, the project implements a modular software architecture with a FastAPI-based web workflow, including file upload, parameter validation, asynchronous job processing, status tracking, and MIDI download. This integration addresses both algorithmic and engineering requirements for practical use. The study is grounded in historical and artistic context, from player pianos to modern Disklavier systems, and is conceptually connected to speech-like mechanical piano practices such as Peter Ablinger’s Deus Cantando. Results demonstrate the feasibility of transforming monophonic speech into expressive symbolic piano performance, while highlighting current limitations related to transcription precision, non-stationary speech complexity, and future real-time.
2025
This thesis presents the design and implementation of a speech-to-MIDI framework for expressive automatic piano playback on a Yamaha Disklavier. The work investigates how human voice signals can be transformed into symbolic musical events through Digital Signal Processing (DSP), then reproduced acoustically through electromechanical piano actuation. The proposed system combines Short-Time Fourier Transform (STFT)-based time-frequency analysis, spectral peak detection, frequency-to-MIDI mapping, and velocity estimation to generate musically interpretable MIDI note streams from speech recordings. In addition to the core conversion pipeline, the project implements a modular software architecture with a FastAPI-based web workflow, including file upload, parameter validation, asynchronous job processing, status tracking, and MIDI download. This integration addresses both algorithmic and engineering requirements for practical use. The study is grounded in historical and artistic context, from player pianos to modern Disklavier systems, and is conceptually connected to speech-like mechanical piano practices such as Peter Ablinger’s Deus Cantando. Results demonstrate the feasibility of transforming monophonic speech into expressive symbolic piano performance, while highlighting current limitations related to transcription precision, non-stationary speech complexity, and future real-time.
Speech-to-MIDI
Yamaha Disklavier
STFT
Pitch Detection
Spectral Analysis
File in questo prodotto:
File Dimensione Formato  
FATAHI_GHAZAL.pdf

accesso aperto

Dimensione 1.17 MB
Formato Adobe PDF
1.17 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/109381