Managing and validating technical or legal documents requires time and precision. Standard manual review procedures are often slow and inherently prone to oversights. With this in mind, this paper describes the development of the "AI Document Consistency Platform," a web app designed to automate quality control of text documents. Functionally, the system allows for the import of documents in various formats (PDF, Word, Markdown) and subjects them to multi-level analysis. Specifically, the platform verifies content consistency, the presence of mandatory sections, compliance with internal policies, and the correct length of sections. The project's architecture requires a hybrid approach: structural and deterministic checks are managed through algorithmic (rule-based) logic, while semantic interpretation of the text is delegated to agents based on Large Language Models, specifically integrating Anthropic's Claude models. Once the analysis phase is complete, the application provides a structured report that highlights any inaccuracies found and then suggests corrective actions. The practical goal of the project is to demonstrate how artificial intelligence can dramatically streamline review workflows while ensuring high levels of standardization. Regarding the software architecture, the adopted technology stack includes a user interface built in React, a backend based on the NestJS framework, and persistence of analyzed documents and validation metrics is handled by the non-relational MongoDB database. Files are stored in an Amazon S3 bucket.
Gestire e validare documenti tecnici o legali richiede tempo e precisione. Le normali procedure di revisione manuale sono spesso lente e intrinsecamente soggette a sviste. Partendo da questa problematica, l'elaborato descrive lo sviluppo della "AI Document Consistency Platform", una web app progettata per automatizzare il controllo di qualità sui documenti testuali. Dal punto di vista funzionale, il sistema permette di importare documenti in diversi formati (PDF, Word, Markdown) e li sottopone a un'analisi a più livelli. Nello specifico, la piattaforma verifica la coerenza contenutistica, la presenza delle sezioni obbligatorie, le conformità rispetto a policy interne e la corretta lunghezza delle sezioni. L’architettura del progetto richiede un approccio ibrido: i controlli strutturali e deterministici sono gestiti tramite logiche algoritmiche (rule-based), mentre l'interpretazione semantica del testo è delegata ad agenti basati su Large Language Models, integrando nello specifico i modelli Claude di Anthropic. Terminata la fase di analisi, l’applicativo restituisce un report strutturato che evidenzia all'utente eventuali scorrettezze riscontrate, per poi proporne le relative correzioni. Lo scopo pratico del progetto è dimostrare come l'intelligenza artificiale possa alleggerire drasticamente i flussi di revisione, garantendo al contempo un'elevata standardizzazione. Per quanto riguarda l'architettura del software, lo stack tecnologico adottato prevede: un interfaccia utente realizzata in React, il backend basato sul framework NestJS, mentre la persistenza dei documenti analizzati e delle metriche di validazione è affidata al database non relazionale MongoDB, l'archiviazione dei file avviene su un bucket Amazon S3.
AI Document Consistency Platform: Controllo qualità documentale automatizzato
CANAZZA, ANGELA
2025/2026
Abstract
Managing and validating technical or legal documents requires time and precision. Standard manual review procedures are often slow and inherently prone to oversights. With this in mind, this paper describes the development of the "AI Document Consistency Platform," a web app designed to automate quality control of text documents. Functionally, the system allows for the import of documents in various formats (PDF, Word, Markdown) and subjects them to multi-level analysis. Specifically, the platform verifies content consistency, the presence of mandatory sections, compliance with internal policies, and the correct length of sections. The project's architecture requires a hybrid approach: structural and deterministic checks are managed through algorithmic (rule-based) logic, while semantic interpretation of the text is delegated to agents based on Large Language Models, specifically integrating Anthropic's Claude models. Once the analysis phase is complete, the application provides a structured report that highlights any inaccuracies found and then suggests corrective actions. The practical goal of the project is to demonstrate how artificial intelligence can dramatically streamline review workflows while ensuring high levels of standardization. Regarding the software architecture, the adopted technology stack includes a user interface built in React, a backend based on the NestJS framework, and persistence of analyzed documents and validation metrics is handled by the non-relational MongoDB database. Files are stored in an Amazon S3 bucket.| File | Dimensione | Formato | |
|---|---|---|---|
|
thesis.pdf
Accesso riservato
Dimensione
2.9 MB
Formato
Adobe PDF
|
2.9 MB | Adobe PDF |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/111039