AI-generated vulnerability patches require human review

6 points by ubernostrum


ubernostrum

Across six recently-disclosed CVEs, we produced 6,080 patches using two frontier, cyber-capable reasoning models. The average success rate for generating a patch that fully resolved the vulnerability (without materially changing application behavior) was just 26.0%.

And:

LLM-generated patches did not resolve the vulnerability, added a new vulnerability, or both, an average 53.9% of the time.

I also love their acronym: "Fix-Like Artifacts With Embedded Defects" (FLAWED).

olliej

LLVM is getting an endless stream of these garbage LLM PRs and they’re all bad and all wrong and include clearly LLM justifications that people have learned to rewrite to hide that the PR is a slop machine pr.