In recent years, the amount of AI-generated images has exponentially increased, making reliable metrics increasingly important for evaluating the aesthetics and visual quality of synthetic visual content. A generated image may be sharp, colorful and apparently coherent, while still containing geometric inconsistencies, unnatural proportions, implausible spatial relations or subtle artifacts that affect the way humans perceive its quality. On the contrary, an image may be geometrically plausible but aesthetically weak because of poor composition, lighting, color harmony or lack of expressive intent. This thesis investigates aesthetic assessment for AI-generated images by analyzing the capabilities and limitations of current state-of-the-art evaluation models. Fine-Grained Image Aesthetic Assessment is adopted as the baseline model, and its predictions are compared with scores produced by HPSv3 and with judgments produced by a Vision-Language Model. To better approximate human judgment, both VLMs and HPSv3 are employed to perform pairwise comparisons between images, providing both preference decisions and textual explanations. These comparisons are then used to construct a high-quality preference dataset for fine-tuning FG-IAA, allowing the model to better align its predictions with perceptual preferences. Experimental results show that Vision-Language Models can provide meaningful semantic feedback and identify aesthetic aspects that are often overlooked by conventional scoring models. Furthermore, incorporating pairwise preference supervision improves the consistency of aesthetic predictions, highlighting the potential of combining traditional aesthetic assessment models with large multimodal models for more human-aligned image quality evaluation.

In recent years, the amount of AI-generated images has exponentially increased, making reliable metrics increasingly important for evaluating the aesthetics and visual quality of synthetic visual content. A generated image may be sharp, colorful and apparently coherent, while still containing geometric inconsistencies, unnatural proportions, implausible spatial relations or subtle artifacts that affect the way humans perceive its quality. On the contrary, an image may be geometrically plausible but aesthetically weak because of poor composition, lighting, color harmony or lack of expressive intent. This thesis investigates aesthetic assessment for AI-generated images by analyzing the capabilities and limitations of current state-of-the-art evaluation models. Fine-Grained Image Aesthetic Assessment is adopted as the baseline model, and its predictions are compared with scores produced by HPSv3 and with judgments produced by a Vision-Language Model. To better approximate human judgment, both VLMs and HPSv3 are employed to perform pairwise comparisons between images, providing both preference decisions and textual explanations. These comparisons are then used to construct a high-quality preference dataset for fine-tuning FG-IAA, allowing the model to better align its predictions with perceptual preferences. Experimental results show that Vision-Language Models can provide meaningful semantic feedback and identify aesthetic aspects that are often overlooked by conventional scoring models. Furthermore, incorporating pairwise preference supervision improves the consistency of aesthetic predictions, highlighting the potential of combining traditional aesthetic assessment models with large multimodal models for more human-aligned image quality evaluation.

Fine-Grained Aesthetic Assessment of AI-Generated Images Using Vision-Language Models and Human Preferences

FARRUKU, JURI
2025/2026

Abstract

In recent years, the amount of AI-generated images has exponentially increased, making reliable metrics increasingly important for evaluating the aesthetics and visual quality of synthetic visual content. A generated image may be sharp, colorful and apparently coherent, while still containing geometric inconsistencies, unnatural proportions, implausible spatial relations or subtle artifacts that affect the way humans perceive its quality. On the contrary, an image may be geometrically plausible but aesthetically weak because of poor composition, lighting, color harmony or lack of expressive intent. This thesis investigates aesthetic assessment for AI-generated images by analyzing the capabilities and limitations of current state-of-the-art evaluation models. Fine-Grained Image Aesthetic Assessment is adopted as the baseline model, and its predictions are compared with scores produced by HPSv3 and with judgments produced by a Vision-Language Model. To better approximate human judgment, both VLMs and HPSv3 are employed to perform pairwise comparisons between images, providing both preference decisions and textual explanations. These comparisons are then used to construct a high-quality preference dataset for fine-tuning FG-IAA, allowing the model to better align its predictions with perceptual preferences. Experimental results show that Vision-Language Models can provide meaningful semantic feedback and identify aesthetic aspects that are often overlooked by conventional scoring models. Furthermore, incorporating pairwise preference supervision improves the consistency of aesthetic predictions, highlighting the potential of combining traditional aesthetic assessment models with large multimodal models for more human-aligned image quality evaluation.
2025
Fine-Grained Aesthetic Assessment of AI-Generated Images Using Vision-Language Models and Human Preferences
In recent years, the amount of AI-generated images has exponentially increased, making reliable metrics increasingly important for evaluating the aesthetics and visual quality of synthetic visual content. A generated image may be sharp, colorful and apparently coherent, while still containing geometric inconsistencies, unnatural proportions, implausible spatial relations or subtle artifacts that affect the way humans perceive its quality. On the contrary, an image may be geometrically plausible but aesthetically weak because of poor composition, lighting, color harmony or lack of expressive intent. This thesis investigates aesthetic assessment for AI-generated images by analyzing the capabilities and limitations of current state-of-the-art evaluation models. Fine-Grained Image Aesthetic Assessment is adopted as the baseline model, and its predictions are compared with scores produced by HPSv3 and with judgments produced by a Vision-Language Model. To better approximate human judgment, both VLMs and HPSv3 are employed to perform pairwise comparisons between images, providing both preference decisions and textual explanations. These comparisons are then used to construct a high-quality preference dataset for fine-tuning FG-IAA, allowing the model to better align its predictions with perceptual preferences. Experimental results show that Vision-Language Models can provide meaningful semantic feedback and identify aesthetic aspects that are often overlooked by conventional scoring models. Furthermore, incorporating pairwise preference supervision improves the consistency of aesthetic predictions, highlighting the potential of combining traditional aesthetic assessment models with large multimodal models for more human-aligned image quality evaluation.
Image Aesthetics
Human Preferences
Preference Learning
File in questo prodotto:
File Dimensione Formato  
Farruku_Juri.pdf

accesso aperto

Dimensione 16.63 MB
Formato Adobe PDF
16.63 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/112993