The discovery of neutrino oscillations established that flavor eigenstates (νe,νμ,ντ) are linear superpositions of mass eigenstates (ν1,ν2,ν3). Whether the overall mass hierarchy follows the Normal Ordering (NO), m1 < m2 < m3, or the Inverted Ordering (IO), m3 < m1 < m2, remains an important open question. The Jiangmen Underground Neutrino Observatory (JUNO) is a 20 kton liquid scintillator detector primarily designed to resolve the Neutrino Mass Ordering (NMO) problem. JUNO detects reactor electron antineutrinos via the inverse beta-decay reaction; its data-taking started in August 2025. Determining the NMO relies on observing a fine-grained interference pattern in the antineutrino energy spectrum. Achieving the sensitivity required to distinguish between the NO and IO expected spectra depends critically on reaching an unprecedented energy resolution of 3% at 1 MeV. This resolution is degraded by the spatial non-uniformity of the detector response, requiring mitigation through calibration. Indeed, identical energy deposition at different points yields different signals across the 17596 PMTs due to complex optical processes like photon attenuation, scattering and geometric shadowing. For this reason, a high-precision energy reconstruction depends on an accurate vertex reconstruction. The official likelihood-based algorithm in use, OMILREC, retains a position-dependent vertex bias of up to 13 cm near the detector boundaries, which directly degrades the energy resolution and restricts the fiducial volume available for physics analyses. This thesis addresses this limitation by developing and evaluating Machine Learning (ML) models for the vertex reconstruction task. Moving beyond previous studies limited to Monte Carlo (MC) simulations, this work provides a first comprehensive assessment of the performance achievable using currently available calibration data. The analysis focuses on a one-dimensional reconstruction, framed as a supervised regression task. Models are trained utilizing experimental data from ⁶⁸Ge and Am-C radioactive sources, deployed at known discrete positions to provide target labels along the detector's central vertical axis. Four architectures were developed and benchmarked against OMILREC: a Fully Connected Neural Network operating on a set of 133 macroscopic aggregated features, and three architectures processing the full, PMT-level hit sequences. These include a Convolutional Neural Network operating on image-like projections of the PMT hit pattern, and two attention-based architectures processing the variable-length set of fired PMT hits: a standard Transformer-Encoder, and the Point Transformer V3 (PTv3). Extensive hyperparameter optimization and comparison among different architectures reveal that the current performance limits are mostly dictated by data insufficiency—specifically, the sparse spatial sampling and the highly unbalanced energy distribution of the discrete calibration sources—rather than by architectural constraints. With the available dataset, no single architecture is unambiguously superior across all performance metrics, nor can it be concluded that there is a clear advantage in processing the full PMT-level information over a well-engineered set of aggregated features. Among the attention-based models, PTv3 substantially mitigates the computational bottleneck inherent to the Transformer-Encoder, proving to be the most viable candidate of the two for a future full-scale deployment. Ultimately, this work demonstrates that training the ML models exclusively on the discretely sampled calibration data currently available cannot guarantee a reliable and uniform reconstruction across the full detector volume. None of the developed models is yet ready to replace OMILREC. The integration of MC simulated data, once available, is identified as essential for future developments.
Vertex Reconstruction in the JUNO Experiment via Machine Learning Techniques
CAVALLIN, JONATHAN
2025/2026
Abstract
The discovery of neutrino oscillations established that flavor eigenstates (νe,νμ,ντ) are linear superpositions of mass eigenstates (ν1,ν2,ν3). Whether the overall mass hierarchy follows the Normal Ordering (NO), m1 < m2 < m3, or the Inverted Ordering (IO), m3 < m1 < m2, remains an important open question. The Jiangmen Underground Neutrino Observatory (JUNO) is a 20 kton liquid scintillator detector primarily designed to resolve the Neutrino Mass Ordering (NMO) problem. JUNO detects reactor electron antineutrinos via the inverse beta-decay reaction; its data-taking started in August 2025. Determining the NMO relies on observing a fine-grained interference pattern in the antineutrino energy spectrum. Achieving the sensitivity required to distinguish between the NO and IO expected spectra depends critically on reaching an unprecedented energy resolution of 3% at 1 MeV. This resolution is degraded by the spatial non-uniformity of the detector response, requiring mitigation through calibration. Indeed, identical energy deposition at different points yields different signals across the 17596 PMTs due to complex optical processes like photon attenuation, scattering and geometric shadowing. For this reason, a high-precision energy reconstruction depends on an accurate vertex reconstruction. The official likelihood-based algorithm in use, OMILREC, retains a position-dependent vertex bias of up to 13 cm near the detector boundaries, which directly degrades the energy resolution and restricts the fiducial volume available for physics analyses. This thesis addresses this limitation by developing and evaluating Machine Learning (ML) models for the vertex reconstruction task. Moving beyond previous studies limited to Monte Carlo (MC) simulations, this work provides a first comprehensive assessment of the performance achievable using currently available calibration data. The analysis focuses on a one-dimensional reconstruction, framed as a supervised regression task. Models are trained utilizing experimental data from ⁶⁸Ge and Am-C radioactive sources, deployed at known discrete positions to provide target labels along the detector's central vertical axis. Four architectures were developed and benchmarked against OMILREC: a Fully Connected Neural Network operating on a set of 133 macroscopic aggregated features, and three architectures processing the full, PMT-level hit sequences. These include a Convolutional Neural Network operating on image-like projections of the PMT hit pattern, and two attention-based architectures processing the variable-length set of fired PMT hits: a standard Transformer-Encoder, and the Point Transformer V3 (PTv3). Extensive hyperparameter optimization and comparison among different architectures reveal that the current performance limits are mostly dictated by data insufficiency—specifically, the sparse spatial sampling and the highly unbalanced energy distribution of the discrete calibration sources—rather than by architectural constraints. With the available dataset, no single architecture is unambiguously superior across all performance metrics, nor can it be concluded that there is a clear advantage in processing the full PMT-level information over a well-engineered set of aggregated features. Among the attention-based models, PTv3 substantially mitigates the computational bottleneck inherent to the Transformer-Encoder, proving to be the most viable candidate of the two for a future full-scale deployment. Ultimately, this work demonstrates that training the ML models exclusively on the discretely sampled calibration data currently available cannot guarantee a reliable and uniform reconstruction across the full detector volume. None of the developed models is yet ready to replace OMILREC. The integration of MC simulated data, once available, is identified as essential for future developments.| File | Dimensione | Formato | |
|---|---|---|---|
|
Cavallin_Jonathan.pdf
Accesso riservato
Dimensione
9.32 MB
Formato
Adobe PDF
|
9.32 MB | Adobe PDF |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/113153