From Description to Score: Can LLMs Quantify Vulnerabilities? -- The 41st ACM/SIGAPP Symposium On Applied Computing (SAC 2026) Academic Article uri icon

Abstract

  • Manual vulnerability scoring, such as assigning Common Vulnera-bility Scoring System (CVSS) scores, is a resource-intensive processthat is often influenced by subjective interpretation. This studyinvestigates the potential of general-purpose large language mod-els (LLMs), namely ChatGPT, Llama, Grok, DeepSeek, and Gemini,to automate this process by analyzing over 31,000 recent Com-mon Vulnerabilities and Exposures (CVE) entries. The results showthat LLMs substantially outperform the baseline on certain met-rics (e.g., Availability Impact), while offering more modest gainson others (e.g., Attack Complexity). Moreover, model performancevaries across both LLM families and individual CVSS metrics, withChatGPT-5 attaining the highest precision. Our analysis reveals thatLLMs tend to misclassify many of the same CVEs, and ensemble-based meta-classifiers only marginally improve performance. Fur-ther examination shows that CVE descriptions often lack criticalcontext or contain ambiguous phrasing, which contributes to sys-tematic misclassifications. These findings underscore the impor-tance of enhancing vulnerability descriptions and incorporatingricher contextual details to support more reliable automated rea-soning and alleviate the growing backlog of CVEs awaiting triage.

Publication Date

  • 2026-06-01

Published In

  • ACM  Journal