Security advisories are the primary way vendors communicate information about vulnerabili- ties affecting their products to the analysts and organizations that must act on them, yet they are published in unstructured, human-readable formats that vary widely in structure, terminol- ogy, and level of detail. Although the OASIS Common Security Advisory Framework (CSAF) provides a machine-readable standard for structuring this information, most vendors publish primarily in free text, leaving analysts to manually and often inconsistently extract critical de- tails. Before this process can be meaningfully automated, however, there must first be a way to measure, systematically, how this information is communicated. This thesis develops such a method, evaluating the presence and clarity with which vulnerability information is commu- nicated in security advisories, and investigates whether Large Language Models (LLMs) can support or scale this evaluation. To operationalize this evaluation, we developed a codebook that scores advisories on two dimensions—binary Presence and four-point ordinal Clarity—across three categories of infor- mation: Advisory, Product, and Vulnerability data. Using CSAF v2.1 as a structural reference, the codebook was derived by systematically analyzing 67 CSAF properties and sub-properties for their relevance to human-readable advisory content and consolidating them into 14 code- book properties. After iterative pilot refinement, the codebook was validated in a human eval- uation study in which five raters with background in cyber security independently coded six real-world security advisories from diverse vendors. The same security advisories and code- book were then given to three state-of-the-art LLMs under a controlled, zero-shot, structured- output, enabling a direct comparison between human and automated evaluation. Presence judgements reached almost perfect agreement among human raters (mean Fleiss’ κ = 0.859, Gwet’s AC1 = 0.954), showing that the 14 properties are clearly enough defined to be identified consistently. Clarity judgements reached substantial but lower agreement (pair- wise mean Gwet’s AC2 = 0.789; advisory-level mean AC2 = 0.7243), confirming that as- sessing communication quality is inherently more subjective than simply detecting whether information is present. Beyond validating the coding scheme itself, applying the codebook sur- faced real differences in advisory quality: Vulnerability Preconditions, for example, was rated uniformly poor by every human and LLM evaluator in one advisory yet unanimously Very Good in another, showing that the codebook captures genuine gaps in vendor security advi- sories rather than only disagreement between coders. Against the human agreement benchmark, LLMs closely approximated human Presence judgements (mean κ = 0.908, compared with a human–human mean of κ = 0.922), confirming that current models can reproduce this judgement reliably at scale. On Clarity, aggre- gate human–LLM agreement (mean AC2 = 0.781) was similarly close to the human–human benchmark (0.789), but individual models’ agreement ranged more widely, from AC2 = 0.743 to 0.825, and showed systematic, model-specific differences in how they used the rating scale meaning aggregate agreement alone is not sufficient grounds to trust automated Clarity evalua- tion. Overall, the codebook proved to be a reliable and practically useful instrument for system- atically evaluating the presence and clarity of vulnerability information in security advisories, establishing a foundation for scalable, AI-assisted assessment. LLMs approach human perfor- mance closely enough on Presence to support large-scale automated screening, while Clarity judgements should, for now, remain primarily a human task within a hybrid evaluation work- flow.
Security advisories are the primary way vendors communicate information about vulnerabili- ties affecting their products to the analysts and organizations that must act on them, yet they are published in unstructured, human-readable formats that vary widely in structure, terminol- ogy, and level of detail. Although the OASIS Common Security Advisory Framework (CSAF) provides a machine-readable standard for structuring this information, most vendors publish primarily in free text, leaving analysts to manually and often inconsistently extract critical de- tails. Before this process can be meaningfully automated, however, there must first be a way to measure, systematically, how this information is communicated. This thesis develops such a method, evaluating the presence and clarity with which vulnerability information is commu- nicated in security advisories, and investigates whether Large Language Models (LLMs) can support or scale this evaluation. To operationalize this evaluation, we developed a codebook that scores advisories on two dimensions—binary Presence and four-point ordinal Clarity—across three categories of infor- mation: Advisory, Product, and Vulnerability data. Using CSAF v2.1 as a structural reference, the codebook was derived by systematically analyzing 67 CSAF properties and sub-properties for their relevance to human-readable advisory content and consolidating them into 14 code- book properties. After iterative pilot refinement, the codebook was validated in a human eval- uation study in which five raters with background in cyber security independently coded six real-world security advisories from diverse vendors. The same security advisories and code- book were then given to three state-of-the-art LLMs under a controlled, zero-shot, structured- output, enabling a direct comparison between human and automated evaluation. Presence judgements reached almost perfect agreement among human raters (mean Fleiss’ κ = 0.859, Gwet’s AC1 = 0.954), showing that the 14 properties are clearly enough defined to be identified consistently. Clarity judgements reached substantial but lower agreement (pair- wise mean Gwet’s AC2 = 0.789; advisory-level mean AC2 = 0.7243), confirming that as- sessing communication quality is inherently more subjective than simply detecting whether information is present. Beyond validating the coding scheme itself, applying the codebook sur- faced real differences in advisory quality: Vulnerability Preconditions, for example, was rated uniformly poor by every human and LLM evaluator in one advisory yet unanimously Very Good in another, showing that the codebook captures genuine gaps in vendor security advi- sories rather than only disagreement between coders. Against the human agreement benchmark, LLMs closely approximated human Presence judgements (mean κ = 0.908, compared with a human–human mean of κ = 0.922), confirming that current models can reproduce this judgement reliably at scale. On Clarity, aggre- gate human–LLM agreement (mean AC2 = 0.781) was similarly close to the human–human benchmark (0.789), but individual models’ agreement ranged more widely, from AC2 = 0.743 to 0.825, and showed systematic, model-specific differences in how they used the rating scale meaning aggregate agreement alone is not sufficient grounds to trust automated Clarity evalua- tion. Overall, the codebook proved to be a reliable and practically useful instrument for system- atically evaluating the presence and clarity of vulnerability information in security advisories, establishing a foundation for scalable, AI-assisted assessment. LLMs approach human perfor- mance closely enough on Presence to support large-scale automated screening, while Clarity judgements should, for now, remain primarily a human task within a hybrid evaluation work- flow.
From Humans to LLMs: A Framework for Evaluating Human-Readable Security Advisories
MUSTAFA, ZIYAD O M
2025/2026
Abstract
Security advisories are the primary way vendors communicate information about vulnerabili- ties affecting their products to the analysts and organizations that must act on them, yet they are published in unstructured, human-readable formats that vary widely in structure, terminol- ogy, and level of detail. Although the OASIS Common Security Advisory Framework (CSAF) provides a machine-readable standard for structuring this information, most vendors publish primarily in free text, leaving analysts to manually and often inconsistently extract critical de- tails. Before this process can be meaningfully automated, however, there must first be a way to measure, systematically, how this information is communicated. This thesis develops such a method, evaluating the presence and clarity with which vulnerability information is commu- nicated in security advisories, and investigates whether Large Language Models (LLMs) can support or scale this evaluation. To operationalize this evaluation, we developed a codebook that scores advisories on two dimensions—binary Presence and four-point ordinal Clarity—across three categories of infor- mation: Advisory, Product, and Vulnerability data. Using CSAF v2.1 as a structural reference, the codebook was derived by systematically analyzing 67 CSAF properties and sub-properties for their relevance to human-readable advisory content and consolidating them into 14 code- book properties. After iterative pilot refinement, the codebook was validated in a human eval- uation study in which five raters with background in cyber security independently coded six real-world security advisories from diverse vendors. The same security advisories and code- book were then given to three state-of-the-art LLMs under a controlled, zero-shot, structured- output, enabling a direct comparison between human and automated evaluation. Presence judgements reached almost perfect agreement among human raters (mean Fleiss’ κ = 0.859, Gwet’s AC1 = 0.954), showing that the 14 properties are clearly enough defined to be identified consistently. Clarity judgements reached substantial but lower agreement (pair- wise mean Gwet’s AC2 = 0.789; advisory-level mean AC2 = 0.7243), confirming that as- sessing communication quality is inherently more subjective than simply detecting whether information is present. Beyond validating the coding scheme itself, applying the codebook sur- faced real differences in advisory quality: Vulnerability Preconditions, for example, was rated uniformly poor by every human and LLM evaluator in one advisory yet unanimously Very Good in another, showing that the codebook captures genuine gaps in vendor security advi- sories rather than only disagreement between coders. Against the human agreement benchmark, LLMs closely approximated human Presence judgements (mean κ = 0.908, compared with a human–human mean of κ = 0.922), confirming that current models can reproduce this judgement reliably at scale. On Clarity, aggre- gate human–LLM agreement (mean AC2 = 0.781) was similarly close to the human–human benchmark (0.789), but individual models’ agreement ranged more widely, from AC2 = 0.743 to 0.825, and showed systematic, model-specific differences in how they used the rating scale meaning aggregate agreement alone is not sufficient grounds to trust automated Clarity evalua- tion. Overall, the codebook proved to be a reliable and practically useful instrument for system- atically evaluating the presence and clarity of vulnerability information in security advisories, establishing a foundation for scalable, AI-assisted assessment. LLMs approach human perfor- mance closely enough on Presence to support large-scale automated screening, while Clarity judgements should, for now, remain primarily a human task within a hybrid evaluation work- flow.| File | Dimensione | Formato | |
|---|---|---|---|
|
MUSTAFA_ZIYAD.pdf
accesso aperto
Dimensione
1.44 MB
Formato
Adobe PDF
|
1.44 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/115864