The development of quantum computers and the associated threats to currently deployed cryptographic mechanisms increase the importance of identifying cryptography present in software. Knowledge of the algorithms and other cryptographic assets in use is important for planning the migration to post-quantum cryptography. A Cryptography Bill of Materials (CBOM) provides a structured representation of information about cryptographic assets used within a software system. This thesis investigates the potential of large language models (LLMs) for detecting cryptographic assets in source code and analyses the capabilities and limitations of an existing approach based on static analysis. An LLM-based approach was designed and implemented using a locally deployed model through Ollama. CBOMKit was used as the baseline static analysis tool. The results of each approach were evaluated against manually prepared ground truth. The evaluation considered detection performance, including the correctness and completeness of the identified cryptographic assets. It also examined the reliability of the results generated by the LLM and the influence of code complexity and available context on detection performance. The results show that the LLM-based approach effectively identified cryptographic assets when the information required for their identification was locally available in the analysed code. In more complex cases, where the relevant context was distributed across the code, detection performance was lower. The model also produced consistent results across repeated executions under the applied configuration. The results indicate the potential of LLMs for detecting cryptographic assets and supporting the generation of CBOMs.
The development of quantum computers and the associated threats to currently deployed cryptographic mechanisms increase the importance of identifying cryptography present in software. Knowledge of the algorithms and other cryptographic assets in use is important for planning the migration to post-quantum cryptography. A Cryptography Bill of Materials (CBOM) provides a structured representation of information about cryptographic assets used within a software system. This thesis investigates the potential of large language models (LLMs) for detecting cryptographic assets in source code and analyses the capabilities and limitations of an existing approach based on static analysis. An LLM-based approach was designed and implemented using a locally deployed model through Ollama. CBOMKit was used as the baseline static analysis tool. The results of each approach were evaluated against manually prepared ground truth. The evaluation considered detection performance, including the correctness and completeness of the identified cryptographic assets. It also examined the reliability of the results generated by the LLM and the influence of code complexity and available context on detection performance. The results show that the LLM-based approach effectively identified cryptographic assets when the information required for their identification was locally available in the analysed code. In more complex cases, where the relevant context was distributed across the code, detection performance was lower. The model also produced consistent results across repeated executions under the applied configuration. The results indicate the potential of LLMs for detecting cryptographic assets and supporting the generation of CBOMs.
Evaluation of LLMs for the detection of cryptographic assets and the generation of a cryptography bill of materials
TRUTY, ADAM ROBERT
2025/2026
Abstract
The development of quantum computers and the associated threats to currently deployed cryptographic mechanisms increase the importance of identifying cryptography present in software. Knowledge of the algorithms and other cryptographic assets in use is important for planning the migration to post-quantum cryptography. A Cryptography Bill of Materials (CBOM) provides a structured representation of information about cryptographic assets used within a software system. This thesis investigates the potential of large language models (LLMs) for detecting cryptographic assets in source code and analyses the capabilities and limitations of an existing approach based on static analysis. An LLM-based approach was designed and implemented using a locally deployed model through Ollama. CBOMKit was used as the baseline static analysis tool. The results of each approach were evaluated against manually prepared ground truth. The evaluation considered detection performance, including the correctness and completeness of the identified cryptographic assets. It also examined the reliability of the results generated by the LLM and the influence of code complexity and available context on detection performance. The results show that the LLM-based approach effectively identified cryptographic assets when the information required for their identification was locally available in the analysed code. In more complex cases, where the relevant context was distributed across the code, detection performance was lower. The model also produced consistent results across repeated executions under the applied configuration. The results indicate the potential of LLMs for detecting cryptographic assets and supporting the generation of CBOMs.| File | Dimensione | Formato | |
|---|---|---|---|
|
Truty_Adam.pdf
Accesso riservato
Dimensione
874.11 kB
Formato
Adobe PDF
|
874.11 kB | Adobe PDF |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/116325