The Era of Software Quality, or the Era of Ostriches?
24 points by calvin
24 points by calvin
Rust will indeed eliminate most memory safety issues (except in unsafe blocks), and you can reasonably expect a Rust project to have an order of magnitude fewer vulnerabilities than a comparable project written in C or C++ or Vala.
Currently the best solution is to not use programming language package managers, but GNOME’s Rust code depends heavily on Cargo. Accordingly, I recommend against using Rust for writing GNOME software.
In my opinion the response should be to pin versions of the dependencies or fork and maintain them in the GNOME project, rather than avoiding a programming language the author claims reduces vulnerabilities by an order of magnitude.
The author seems to conflate “using programming language package managers” with “having uncontrolled dependencies”, which I don’t think is valid. I’m pretty sure you could configure a Rust build with controlled dependencies. (And a C++ build with uncontrolled ones.)
Rust has no inherent coupling to Cargo. Rust codebases can be built with make, cmake, or any other build system by directly invoking rustc in the same manner one would directly invoke gcc. For Rust projects whose build system is Cargo, the cargo vendor subcommand conveniently produces a vendored tree of dependencies amenable to version control.
Right. I suspect what the author meant was more like “you can’t write a useful Rust program without using lots of external packages”. A comment on the ecosystem, not the language.
Of course, you could also say “you can’t write a useful C++ program without reinventing a bunch of wheels insecurely”.
Indeed, there are absolutely more ways to help deal with malicious dependencies than memory safety. It seems insane to me to write a blog post on security issues and then move past an easy solution for a huge chunk of them because of this.
Having not used Cargo recently but having worked with Python, Node, and Elixir a fair bit… how many dependencies are we talking about here? From working with those other ecosystems I am assuming that we’re talking at least hundreds of transitive dependencies.
I'm a big Rust enthusiast but I think that metric misses the point for this particular discussion. For supply chain vulnerabilities, I think that number of unique packages (or perhaps more accurate would be number of unique package authors) is the most important metric, because if you are an outsider trying to perform a supply chain attack, you are looking to compromise any single package in the dependency tree of your target.
I think the LoC or bytes metrics are definitely relevant when discussing bloat of a library or application, and I get a bit irked when I see the number of Cargo packages being pulled in as an argument about that, but in the supply chain case I do think it is unfortunately quite relevant.
I don't have a good solution to this besides automated scanning of crates published, as well as a cooldown period in Cargo (or one of the wrappers or crate mirrors that provides this), but it still feels incomplete to me.
Why is Vala listed along C and C++? I thought it had automatic memory management (with the option to opt out of it)
Vala compiles to C. It's got memory management constructs, but they aren't checked like in Rust.
void main () {
unowned string x; // this is the equivalent of a reference with a lifetime from outside the function...
{
string y = "heap-allocated".up ();
x = y; // ... but this just generates a pointer copy
}
print ("%s\n", x); // use after free
}
More trivially, arrays are just not bounds-checked:
void main () {
int[] a = { 1, 2, 3 };
a[4] = 0; // buffer overflow
}
Vala compiles to C
So do many Lisp implementations (Chicken, ECL, etc). That doesn't mean they are memory unsafe.
More trivially, arrays are just not bounds-checked:
Ok, didn't know that about Vala. Sounds like it has a horrible implementation. Thanks for the info!
I’m trying to contribute to clippy now and the situation looks bleak for maintainers.
Previously the signal was: could make a PR. Things like docs and tests were "wow this person is good" and a description that linked to the issue it closes and actually explained that it did was a 10/10.
Mantainers have always relied on signals to determine where to invest their limited time. When they do invest it, it was clear their empathy is met and reciprocated. Now those signals are gone and there's not a good way to distinguish between "I spent days and hours meticulously promoting fans getting reviews and iterating combined with my own human iteration" and "told Claude to get me internet points to boost my resume in a Ralph loop". There's not even a guarantee the person on the other side is in fact a person.
The rust/clippy policy is very thoughtful and clearly spells out how to use LLMs with empathy for the eventual maintainer/reviewer. I ran afoul of it by some degrees and even though I'm in the "meticulous code" camp of LLM users cannot genuinely say code produced by the machine isn't in the final product. I can though, and will. So closed it and will rewrite it.
I think LLM scanning is good. But LLM issues and PRs are not. We need a human who has the skills and empathy for reviewers to make it easy enough for them. The challenge has never been the work, the challenge has always been doing the work in a way that is reviewable, maintainable, and sustainable. We need to find community standards and guidelines to reflect that reality.
I think the empathy gap is challenging here. At this point, people have a different sense of how important code is. Some people have gotten really used to shipping products quickly. Friction is anathema to them.
For some work projects, I found myself writing detailed replies to some PRs only to have them answered immediately. That indicated to me that the person didn't even read what I wrote. They just told Claude, "hey, there is a reply. Go handle it."
I fear that approaching an overwhelming onslaught of automated contributions with empathy is just fuel for burn out.
Lots of these effort signals is something I've been queue for a while, eg see https://blog.happyfellow.dev/simulacrum-of-knowledge-work/
There is zero hope of maintaining quality software in 2026 without AI vulnerability scanning. Any claims to the contrary are unserious and delusional.
You're kidding me, right? What about, I dunno, not using those unsafe languages? What about just using a vector sometimes instead of utilizing pointers 24/7?
This is just a bad faith argument to lionize LLM usage.
Even if you don't like Rust, I guarantee you that you could write a language that actually suits your needs and compile it to LLVM. It's much easier than banging your head against the wall, i.e. keep writing C/C++/Vala and hope it doesn't blow up.
Ok but what if it’s possible to still write bugs in rust?
Most bugs don’t really matter; what matters are vulnerabilities. And Rust is really good at preventing memory corruption bugs.
There are so many vulnerabilities that can be present without memory corruption. Not checking for an access control is perfectly memory safe
Unsafe languages need to deal with all the same high-level bugs (like access control or SQL injection) in addition to memory safety bugs. C is not trading memory unsafety for some other safety, it's strictly more unsafe. Rust does a lot to prevent logic bugs, e.g. error handling can be enforced by the compiler, while C can't even decide whether returning 0 means OK or false.
Memory unsafety vulnerabilities are exceptional in their unbound scope and severity.
Without arbitrary code execution, the danger is roughly limited to what the code is supposed to do, and plenty of programs have lots of boring code that doesn't touch anything sensitive. Even if you're implementing access control, the worst case without RCE is failing open on that endpoint, rather than enabling installation of a rootkit or such.
A memory-safe JSON parser can at worst give you garbage data or heat your CPU or eat your RAM, but it won't run arbitrary attacker-supplied code. If you're implementing a widget that queries the current weather, the the worst case is not getting the weather info. In C, the worst case is getting the whole system pwned through a weather info widget. It's a massive difference in the attack surface.
That LLMs have been finding bugs is more so an effect of the tech industry not taking security seriously than them being amazing at it. It also doesn't hurt that GenAI companies are pumping millions into making their models look as amazing as possible. If you hire a team of pentesters for a few million bucks, I'm sure you'll get way more (both quality and quantity) than whatever the LLMs regurgitated.
If planet-burning slop machines are what it takes for the tech industry to start taking security seriously... heh. Though, that being said, throwing slop-machines at it doesn't mean they're taking it seriously, so I don't even think it had that effect. Either way, all we can do is object to this extremely unethical technology.
If you hire a team of pentesters for a few million bucks, I'm sure you'll get way more (both quality and quantity) than whatever the LLMs regurgitated.
How about hiring pentesters for the ~$100 it takes to run a comprehensive security scan with a frontier model?
Ah, yes, ignoring all of the negative externalities (like pollution), massive capital investment in data centers, LLM companies not turning a profit, etc. we can pretend the slop machines are really cheap!
If a drug dealer offers me the first hit for free, that's free drugs, that's fantastic, I love free drugs.
Isn't it pretty clear though that for something like a security scan or similar that even the externalities included price far cheaper than the alternative? Whether that's a good thing or not can still be an open question but I'd be interested in hearing the argument that LLMs are only cost effective for finding vulnerabilities because tokens are subsidized/should cost more for some other reason
My silly pet project (the pi searcher) is in Rust, and I still got some useful feedback about a potential XSS and a potential high resource use DoS path on my code from an LLM. Nothing terrible or RCE-ish but it was still useful.
What about, I dunno, not using those unsafe languages?
This is presumably a problem for the large quantity of existing software written in unsafe languages that is not going to be rewritten overnight (as well as things that need to toe the boundary frequently like JITs).
We’ve spent more than a decade seeing anti-Rust people say “Rust doesn’t prevent all security issues”, with everyone patiently explained “of course not, but Rust is still worth it.”
So it’s wild to see a post that actually seems to imply Rust is an adequate solution to all security problems—no need to use automated vulnerability scans once you’ve used Rust, and no need to try and improve the security of the millions (billions?) of existing lines of C and C++ code in the wild.
The very definition of "Software quality" has shifted due to the adversarial effects of LLMs being widespread, to a degree that we now must further lionize LLMs to mitigate this shift. It feels sad that we have allowed this to happen.
The problems with the quality of our software were there all along. We merely have discovered a cost-effective way to find them (which is a good thing!).
It feels sad that we have allowed this to happen.
Yes, agreed. It is profoundly sad that we allowed the industry (and ourselves) to play ostrich and neglect software quality for such a long time. Thankfully, these times are over.
Thankfully, these times are over.
Are they? For me a natural response to being forced to be honest about the broken state of most systems software would be massive rewrite efforts with better tools (memory-safe language, capability-based security, judicious use of formal methods, etc.). But what I see happening in practice is more "exhaust ourselves fixing all the LLM-found issues and then when the current stream dries out we can claim victory, without otherwise changing our programming approach". And duh, of course, this is more incremental and the upcost front is smaller, and the result is worse, so of course everyone does it.
And duh, of course, this is more incremental and the upcost front is smaller, and the result is worse, so of course everyone does it.
It is entirely well known that in every human field of endeavor there is a tradeoff between costs and benefits, and not every thing that needs doing calls for maximum effort. Unless and until someone successfully implements a post-scarcity economy, it is going to stay that way.
Thus, something that raises the zero-effort floor (even if it doesn't change the ceiling or the amount of effort required to attain the ceiling) is still valuable—maybe even more valuable than advancing the ivory-tower ceiling.
I don't think you do agree with me.
The problems with the quality of our software were there all along.
You could make the same argument concerning any system designed to live in an equilibrium with the environment it was created for. I wouldn't say cities built by the coast are poor quality, or had problems that were there all along, because they do not have flood defences, even if it is now becoming a major problem due to a shifting climate.
I don't mean to say that you are wrong for being glad we are finding vulnerabilities, perhaps it leads to a better place. I hope however you can understand that I feel sad about it, as I feel that the benefits we see from these things are outweighed by the negative externalities of their existence and widespread adoption, in a similar way to how the delicious wine we can now grow in my country is not worth the externalities of the climate crisis.
Although Rust mostly eliminates memory safety risk, any use of Cargo to download dependencies dramatically increases supply chain security risk. The risk of bundling a trojanized dependency arguably — I would even say probably — outweighs the benefit of eliminating memory safety flaws. This problem is inherent to any programming language package manager. Currently the best solution is to not use programming language package managers, but GNOME’s Rust code depends heavily on Cargo. Accordingly, I recommend against using Rust for writing GNOME software.
I totally agree regarding the risk of using Cargo dependencies, but at the end, the reasoning is kinda flawed IMO. Why would GNOME’s official libraries for Rust be compromised? Why GNOME software would use Cargo to fetch something else?
I agree that there is a lot of suspicious libraries on https://crates.io/. But that does not mean it’s entirely insecure. It’s a bit like saying “you shouldn’t use internet at all because there are some trojans in the wild”…
I’m kind of disappointed that LLM review is so effective. No specification, no simulation testing, no structured method.
My guess is that it reveals that LLMs can really pay attention?
I’m kind of disappointed that LLM review is so effective. No specification, no simulation testing, no structured method.
Why is it disappointing to you? By all accounts, the fact that something works better than one expected, with less prerequisites than one expected, should be a welcome news?
I think it’s the bitterness of the bitter lesson. It’s an area where LLMs have already surpassed programmers.
I suppose the interesting question is whether LLMs can find the same bugs as eg generative test suites or not, or if the result sets are at least partially disjoint.
I suppose the interesting question is whether LLMs can find the same bugs as eg generative test suites or not, or if the result sets are at least partially disjoint.
Hm, what exactly is a “generative test suite”? Is this another way to spell “fuzzing”, or something else entirely? Google says this refers to something with LLMs (again), which I’m wary of believing as you seem to contrast the two.
Well it depends on your definition of fuzzing :) but sure, fuzzing, deterministic simulation testing, and the like.