Is it too much to ask devs to use AI to review their hand-crafted code?
8 points by vrypan
8 points by vrypan
You believe coding by hand is the Way and I both admire and appreciate beautiful, hand-crafted code. But I want to use your library, and when I ask my AI to evaluate it for bugs, memory leaks, security issues, it returns a long list of serious ones.
Even if you hate AI-generated code, you can use an AI to identify issues and you can fix them yourself. Think of it as an advanced linter or something.
I think a lot of people (myself included) have objections to AI that are unrelated, or in addition, to its (lack of) capabilities. Even if an LLM might find bugs, they might feel that this is not worth the widespread personal and societal cost of using and training them. Yes, that means that a library might have bugs that an LLM may have helped remove, but individual engineers will have to offset that against the impacts of LLM use and make their own judgment.
Even if an LLM might find bugs, they might feel that this is not worth the widespread personal and societal cost of using and training them.
I don't think LLMs are worth those costs but we have them now anyway.
Even if an LLM might find bugs, they might feel that this is not worth the widespread personal and societal cost of using and training them.
A significant chunk of the total environmental cost of LLM use comes from using high-end frontier models with trillions of parameters. This locks you into a US-based duopoly (Google barely counts right now). And large models also require far more compute and electricity to serve. The whole experience feels increasingly bad.
Small models can be run locally (or on shared infrastructure), and they use less energy than the average toaster (300W during active inference is typical). And far less than an electric over or clothes dryer. Training costs are higher--many small to medium models are in the range of a small town's power cost to train--but much of that is done using China's solar infrastructure, and we could easily get buy with training a few models per year. Compared to anything else industrial civilization does, training a few smaller models is pretty minor.
The other advantage of smaller models is that you can get weights and serve them using open source software. This is still fairly expensive at the moment. You'd need around US$1,800 for a good graphic card (say an AI PRO R9700), and a decent PC to stick it in. Which probably isn't worth it, at least not until RAM priced drop. But I think it's important to point out that running an AI agent doesn't require giving money to a giant duopoly that's already starting to enshittify and raise prices. You could do it in a cabin in the woods running of solar panels if you really wanted to.
Now, none of this is a reason to use LLMs. They are, at best, a mixed blessing. And many common ways of using LLMs lead to both rapid skill atrophy and truly awful software quality once you get above a few thousand lines.
Yeah, those are very good points IMO. Even if societal / environmental harms can be mitigated (though I personally feel ambivalent about this even for local models in terms of their training data), I believe that moderate-to-heavy LLM use is bound to have a pretty negative impact on one's ability to learn and understand. I wrote a bit about that a few months ago.
Yes, it is entirely too much to ask. Entitled consumers of FOSS projects are nothing new, if you're not paying for something, you have no right to make demands.
And let's get real, most reasonably long-lived projects contain technical debt galore. You can point an AI or linter or whatever at it and come up with a long list of stuff to fix. But since it works well enough, such stuff is low-priority. Even in settings where people get paid to do it, it usually never gets fixed. In a FOSS project, you can't really ask someone to spend their free time to do it, so either do it yourself or live with it.
Entitled consumers of FOSS projects are nothing new
Come on my friend. I wrote my first code in the 80s, and I set up my first linux machine from a pack of floppy disks.
That isn't really a response to what you are quoting though. You have been around for a long time. That doesn't make the demand for FOSS projects to always use LLMs so you don't experience bugs less entitled. LLM usage does cost money at the very least. Then there are a multitude of other factors that might make someone not want to use LLMs for their project in the first place.
As it stands your OP does not take any of that in account and is basically a demand aimed at your convience. Which, by extension, does make it somewhat of a entitled demand.
You know that stuff costs money, right?
If that's the only thing holding you back, it's not a big issue - there's regularly a new model being tested in stealth for free. With a harness configured for review all of them should be good enough these days. For example see https://openrouter.ai/models?variant=free&output_modalities=text or https://open-code.ai/en/docs/zen#pricing For open source projects you can also use the free tiers used for data collection (the code was already ingested anyway)
Right, but they don't parade with these free offerings, you have to know they exist. Plus you have to figure out if you trust this company with your PII and credit card and all the input/output it processes. We're talking first time users that probably never had an interaction with any LLM providers.
Also: lots of people go entire lives without a credit card! I feel like credit cards are an incredibly US-centric thing.
Also: I found harnesses took real effort to learn initially. Sure you can run something and type a broad question as prompt, but it's more about: figuring out if you can trust this piece of software and this LLM thing with broad access to your system. (Or wrangling a VM into submission for this possible one-off.)
Arguably, Codex's free tier is the easiest path, because I believe you can see it in action without an account and get good results.
You can use all those free models I mentioned without any payment / credit card.
As for trusting the company with the input/output, the context here seems to be open source projects. They also have all that code. You're not trusting them with anything new.
As for familiarity - this applies to all the tools. Linters, security checkers and similar tools you also have to learn about. And I hope someone learned about the options from my message here. It's up to people with information to spread it in places like this.
What's more, I get it. Can I run it for you and submit the fix? Or is it not acceptable because it's AI-generated?
Most open source C code, with several highly admirable exceptions, has always been riddled with security bugs. C requires near super-human diligence to avoid security vulnerabilities. Some maintainers, like curl's Daniel Stenberg, take those vulnerabilities very seriously, which has led to them doing massive amounts of uncompensated work fixing those vulnerabilities. This has become almost intolerably time-consuming in the post-Mythos era. It's not even that the LLM security reports are necessarily bad anymore. Apparently, even just the well-written reports that include fixes already require massive effort to triage.
But some projects have made it clear that they don't want to know, or that they can't possibly keep up with remediation anyway. And a few projects have extremely strong anti-AI policies and they may not welcome you submitting even hand-verified and hand-written-up security fixes that were originally found with AI assistance.
The authors and maintainers of these insecure tools don't owe anybody anything, of course. But I think it might be better all around for everyone to gradually switch to alternative tools written with more secure technology, at least in some cases.
Thanks for framing it in a less loaded way than I did.
My intended point (which did not land well in this crowd) was: Even if want to write code by hand, and don't trust AI tools, consider using them for reviews, they will probably uncover issues you were not aware of.
In other words, you were never interested in others' opinions to your question. You just wanted people to agree that they should use the tools.
If the PR contains input and/or a test case showing the actual bug, and the code follows the coding convention then maybe I won't outright reject it.
If your software is worth less than 20 bucks a month, then I think its worth reflecting on a bit.
I hate to be the one to have to say it, but that is a very first-world-centric attitude. Just as an example, the median salary in South Africa is something like $500 a month. Unemployment among young people is over 50%.
20 bucks a month doesn't seem like anything on a tech salary, but to someone trying to break into tech from poverty conditions, that's an enormous cost.
(And by most accounts, these costs are still heavily subsidized by the model providers. If you go by the common estimate that token plans are subsidized at a 25:1 rate, 20 bucks a month would balloon into literally the entire median salary in SA (and again, that's for the lucky ones who even have a job in the first place!))
Is it too much to ask devs to use AI to review their hand-crafted code?
yes it is, maintainers of OSS don't have to do anything that you ask of them.
It's all over at "I ask my AI". If you just want to be a meat proxy and don't think for yourself, you should at least understand why you're not wanted. You expect others to do the work you don't care to do.
If you're just that much better at OSS contribution, nobody will ever know that you used AI to find bugs because the quality of the work will leave no way to know.
If you're not that much better and it's pretty obvious that you're throwing AI slop at maintainers and expecting them to maintain a project that's better than AI slop, expect to be shown the door. You're as free as anyone to own a fork, and one should think you'd be eager
Depends.
For professional work at a company that already uses AI? Sure, let the machine loose. I think it's one of the better uses of LLMs to only use them as pre-human review.
For a hobby FOSS project, where the author is not interested in using AI? THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS “AS IS”.
Yes, even a minimal subsidised subscription costs 20 euros and this is near 3/4 of my electricity bill. I'm making 0 euros from open source so math doesn't work for me.
https://lobste.rs/c/bn7bla If the cost is the only issue, then you can get around it easily.
Most open source projects are labours of love. They exist because the maintainers like doing it.
Once it starts feeling like work, chances are high that the project gets abandoned or that the maintainer takes a longer break from it.
AI can be helpful, but it can also uncover false positives. It can discover race conditions in code that will never experience that race condition in practice, or simply misunderstand how something works, or outright make something up.
Verifying and fixing something that an AI spits out doesn’t seem particularly interesting to spend ones unpaid, personal time on. And it’s worth pointing out that people who don’t want AI involvement in the first place, probably aren’t interested in paying for it either.
I disagree. I think most open source projects are labour of necessity and they exist because we needed functionality that didn't exist in acceptable form which is why most of it is not really maintained.
Would they be more maintained with use of AI? Probably not.
I think most open source projects are labour of necessity
I would disagree here. Most open-source projects are maybe used ever by one or two people. Most people publish for "look at what I did ma" or otherwise to share something with others they find interesting.
I've been dealing with a whole lot of issues from deep AI code reviews recently and I am learning so much about both my software and software development in general from tackling them.
I'm finding it very interesting (and quite alarming, too.)
I don't want to breathe your second hand smoke, and I'd be especially pissed if you told me to start doing it myself.
Your approach here is both patronizing and entitled.
I will withhold judgement on the original question of whether or not "it is asking too much", because you are simply "asking" for it poorly.
You believe coding by hand is the Way
That is not how you talk to people you respect and want to open a dialogue with.
You say in another comment that you just "don't understand" the "weird politics" around LLMs, but then you start your post by dismissing the maintainer's informed opinions on how to maintain their own library. It seems like you understand just fine, and just don't want to work with people. People can tell.
I want to use your library
Great, the license is the author's permission to do so!
... and when I ask my AI to evaluate it for bugs, memory leaks, security issues, it returns a long list of serious ones.
Ok still fine. Take the time to understand the issues, filter out useless ones, build meaningful reproducers and tests, and then use the project's preferred channels to report the issues and/or provide changesets to resolve them.
From there, it will likely take a while, since this is open-source and people are busy. You should take on the effort yourself up-front to show good-faith, and the maintainers will likely see the care you put into the work, and appreciate how it makes their work easier while improving the library they pour their free time into.
you can use an AI to identify issues and you can fix them yourself
Ok, no longer fine. This implies you feel entitled to all of the labor we just covered above.
Now, you could argue that the original authors have more context, and are more equipped to do a lot of that labor. You could also argue that the LLM reports can be easier to interpret in the context of individual changesets (small) instead of the whole codebase (large). You could steelman the competent maintainers whose libraries are evidently of a high enough quality for you to want to use them to begin with, and then still want to use them in spite of the long list of issues you believe you've identified.
You don't do any of that, though. It seems to me like you don't even want to put in the effort required to make your case for why others should do more work for you.
You are an entitled user of open-source, there is no other way to describe it.
But I want to use your library, and when I ask my AI to evaluate it for bugs, memory leaks, security issues, it returns a long list of serious ones.
I can't stop you from doing that, but I'm not going to do it myself. Plus, if you report bugs to me that don't actually concretely inhibit your usage, I'm banning you from my issue tracker. Security issues are at least somewhat understandable, but effectively all of these will be fairly nondestructive bugs because of the way I write and package software (I don't write code you run in untrusted contexts, and when I do, I aggressively sandbox and use safe code constructs).
I see finding bugs and vulnerabilities as part of the point of writing software for self-edification. Learning from my mistakes and predicting others is an extremely valuable skill, and since the potential for harm with the code that I write is so, so low, I want to do this on my own. Industry programmers probably have different objectives and if they feel they must use these tools, that is at least somewhat understandable, even if I am repulsed by them and by people who try to convince me to use them.
Are you talking about open source repositories for which you are not the maintainer? In which case yes, any ask is too much, you have no authority or right to be heard in any degree.
If you are a paying customer then you may make requests under that banner.
Someone providing a library, for free, is under no obligation to do anything.
AI isn't unlimited or free. I can't afford a $200/month AI subscription. I have a limited $30/month quota and only so much I can do with it. Authors might have other projects they would prefer to use resources on.
If you have the AI quota, it's probably appreciated if you examine each of those issues yourself, ensure they're real, and submit a well-written bug report. It likely won't be appreciated if you just copy/paste your AI output into the project's issue queue.
But I want to use your library, and when I ask my AI to evaluate it for bugs, memory leaks, security issues, it returns a long list of serious ones.
Not on my library it does not. I did lately get bug reports that looked AI generated, all bogus, all based on misuse of the API, and some suggesting fixes that were even worse than the current situation. Which I reckon is not ideal, I'd rather have an impossible to misuse API, or at least total functions. But this is C we're talking about, there's little choice there.
I'd recommend giving it a go. Since you tempted us with "it does not", I tried it out and there are 2 obviously-correct findings that the library can trivially protect from, even if they're user mistakes, including a basic stack overflow. (And a couple of things unlikely to occur in real apps, but the LLM was happy to generate examples proving they're problem - I reviewed manually)
obviously-correct
Have you ever used their library before? Can you qualitatively say that that code region is buggy in the context of the full program? If both and this offends you, why not make a proper pull request to fix it? Why impose this effort on the developer when these "obviously-correct" findings have not affected any user?
Heh, yes, that's a really well-built library.
However, this is worth your attention given the nature of the project:
Most important finding: the download on the homepage is the vulnerable version. The homepage still links monocypher-4.0.2.tar.gz, but the changelog and bugs page say 4.0.2 and below have a timing leak in EdDSA/Ed25519 signing (compiler-dependent)
I just did a silly "check this". I have no idea how to properly ask the AI to evaluate a project I'm not involved in (esp on a topic my capacity is limited). And I'm not familiar with the code, so I have no idea how to evaluate what Fable reported: https://claude.ai/share/b056adee-c43b-4b76-9017-2a13d7311b76
Maybe all of the findings are trash, maybe one of them is of some value to you. But that's my point, it doesn't hurt using an AI to check the code, does it?
But that's my point, it doesn't hurt using an AI to check the code, does it?
Do you not see the hurt right there? To begin with, the time spent?
And I'm not familiar with the code, so I have no idea how to evaluate what Fable reported
Bingo. Part of the reason we don't want AI-generated findings is that they are a time sink, coming from uninformed drive-bys. Why would we want to spend any time on that? Why should we "talk the tool through" to any real conclusion?
You can point out a stale link without any reference to LLMs. It is clearly an oversight, there is no ambiguity. Nobody would mind, I'm sure.
This would only require you write it in your own words.
If you send "the homepage doesn't link to the latest release (and the one it does link has a known bug)" nobody will bat an eye. If the message is a copy-pasted response from an LLM starting with "Most interesting finding: ..." expect to be asked to either value maintainer time more, or go away.
Hmm, good point on the timing leak, 4.0.2 and below are only marked old instead of vulnerable. Not a huge deal I would say (the primary download is the latest, corrected version), but still worth addressing. Thanks.
For the rest of Claude's report:
The source download came back as a binary blob
It's a tarball. I'm surprised Claude could not decompress it.
Argon2 takes a caller-supplied work_area; everything else is stack or caller buffers. So: no leaks, but Argon2 uses ~4–5 KB of stack on top of the work area (tmp, index_block, hash_area, final_block), which matters on MCUs.
The high stack usage is a good point, and does bother me somewhat. Now microcontrollers definitely are not a good target for Argon2 (it's a memory hard hash that makes heavy use of 64-bit multiply). Maybe I could shave off 1 or 2 KB off the stack, but I won't if it hurt performance in any way.
[wiping]
Spot on as far as I can tell. I don't really care though, the threat model where this actually matters is real sketchy in my opinion.
extended_hash deliberately calls BLAKE2b with overlapping input/output. It's safe because the context absorbs the input before _final writes, but it is aliasing that a future refactor could break.
Actually no, a mere refactor cannot break this, because safe aliasing is a documented guarantee of BLAK2B's public API.
crypto_chacha20_ietf returns a u32 counter; past 256 GiB the counter overflows into the nonce word. Documented limit, not a bug, but easy to hit in a file-encryption tool.
Not sure why Claude mentions this: the 256 GiB limit comes from the primitive itself, there's no avoiding it. It's right, though. That's why I provide alternatives with a 64-bit counter, and recommend those in the manual. I only added this one because of the IETF, and the popular demand that followed.
Field arithmetic is ref10-style signed i32[10]. The library fixed the classic negative-left-shift UB back in 0.5, but this representation is inherently full of "we know the compiler does the right thing" reasoning.
Correct. Instead of writing x << 26, which is UB/nasal demons when x is negative, we write x * (1<<26), and rely on the compiler to simplify that back to a left shift. Another point Claude did not mention, is that right shifts on negative integers is implementation defined. But that's okay, all platforms use an arithmetic shift, that propagates the signed bit. There is no exception that I know of.
Missing validation (by design, but your problem)
Not much I can do. At best I can extend the semantics of a couple primitives so they stay safe even when the input ranges go beyond the bounds of the specs. I hesitate to do this though, because soon people may start relying on it, and create vendor lock in. In favour of my library, sure, but the ability to switch away from Monocypher is actually one of its selling points.
See here for a more detailed rationale. Note that I haven't completely ruled out adding asserts everywhere. But for maximum portability they need to be customisable, and probably off by default... and that last one is likely to defeat the point of asserts to begin with.
Track record
Hmm, Claude is being biased by my own reports there, giving an air of authority where there is actually a conflict of interest. I mean, Claude can't really know if there's stuff I haven't reported on, and I do have some incentive to keep the unfixed stuff quiet, at least in the short term.
Still, I'm fairly impressed overall.
Yes, it’s too much to ask. This has nothing to do with AI actually. You are supposed to be willing to do the work yourself, not expecting someone else to do it. Having AI spit out a laundry list is not the hard part.
LLMs do find the occasional bug, it's true, and sometimes that bug probably would not have been noticed by human reviewers. But the LLMs aren't consistent. Sometimes they demand every single constant be represented by a #define (or the equivalent), sometimes they don't. Sometimes they want excruciating null-checks, or thorough array-length checks, sometimes they don't. They get hung up on code that can't follow "best practices", and is excruciatingly commented as to why and how.
These are but a few examples of many categories they're inconsistent about.
There are social side effects. The fact that The Machine Hath Spoken from On High means that a developer can't effectively argue with them: every other human reviewer has had their opinion framed by the LLM's output. The person using the LLM for review has typically turned their critical evaluation skills off, as well, and isn't going to contribute anything more. LLM-generated reviews are much harder to comprehend than human-written reviews. Sentence structure is convoluted. Vocabulary choice is often fairly exotic. Human limitations and cognitive biases come into play, and the aftereffects aren't pretty.
LLM code reviews seem miraculous the first time they happen, but after a while, experience shows you the problems. They're not worth it on the whole.
Here's an example:
The deeper audit of <redacted> found a core bug reachable through valid Unicode text, despite upstream tests passing in Debug and ReleaseFast. Adoption unchanged is discouraged; the next step is to investigate the bug and assess fixes.
That's a u8 holding values that can potentially be 4 or even 8 bytes.
Am I supposed to let the AI fix it and submit a PR? Will I be accused of AI-slop?
If the maintainer does not want to accept contributions, either with AI or in general, you are supposed to fork it and move on with your life. I maintain forks in a professional capacity for my company for similar reasons.
I have an open source project that gets contributions every once in a while, here’s my perspective.
Opening a PR is only acceptable if:
If 1 is ok, but you skip 2 and 3, you’re just adding work to the maintainers, and most of them will not find that helpful.
Same thing applies to making issues.
Pointing an AI at a repo is a start. The real work is ruling out false positives, coming up with the good solution out of a pool of all solutions, and making the end result understandable to others is where most of the work is.
I am an engineer. And I totally agree "coming up with the good solution out of a pool of all solutions, and making the end result understandable to others is where most of the work is" (that's why I don't mind a computer doing the tedious part).
And I could just vendor the library, fix the bug, and solve my problem.
What bothers me is
a) I need to get into these weird politics, which I don't understand. There is no clear no-AI contribution policy, but it's on codeberg, does this mean something? Do I have to dig into some cultural war just to say, "my friend, this u8 should be u32 or u64 or even better, uoffset, thank you for maintaining this"?
b) Other devs using this library introduce a bug to their project that could have been fixed.
Btw, this is one of the simpler cases, in other cases bugs can be deeper and more complex, but also more severe.
Do I have to dig into some cultural war just to say, "my friend, this u8 should be u32 or u64 or even better, uoffset, thank you for maintaining this"?
As long as you were the writer and you’ve actually verified that it is true, then no, not in this specific case (if I’ve interpreted codebergs guidelines correctly)
But in general: yes, at the very minimum you have to read the contribution guidelines before you contribute to a project. Not every project accepts code contributions, human written or otherwise, and it’s not on us to demand differently. Sqlite, for instance, doesn’t accept code from others at all.
Not the point of your comment but sqlite does accept contributions, they just require the copyright to be explicitly waived, have high standards, and may rewrite them.
I think one of the things that separates an engineer from those less experienced is that they essentially see all programming as what you call "weird" politics. The hardest projects are all embedded in a world of real humans with real problems. Engineers don't circumvent that, they work with it as a fundamental basis. Whether it's AI, working in government bureaucracy or building something for your local community, you as an engineer are tasked with holistically attending to all the nuances of the particular problem. There is no pure code that ever escapes that. I like that you are asking questions. There's always something new, AI or not, and reaching out, trying things out, getting it wrong sometimes, it's all part of being a good engineer.
I don't see this as a specifically AI thing, it's just a human thing, they're all weird, the lot of them!
Having contributed often to multiple open source projects, I'd say contributing is a social act. That means you should be mindful of the humans at the other end, including their preferences. They are, after all, the ones who maintain the code and get to choose how they want to work. Even from a pragmatic standpoint, your options are to either: (a) accept whatever criteria maintainers have for what constitutes a valuable contribution; or (b) fork off.
Is that sometimes annoying? Hell yes! It's not fun to have someone review your PR and say you should split the commits some more, use tabs rather than spaces, or sort the imports in a certain way. But if you care about the project as a whole, and not only about your own use case, you put up with it and the project ends up in a better place than it was before.
Has your experience been different when contributing to open source? In my eyes, this all predates the advent of LLMs.
Why are you interested in using open-source? By definition it will require you to navigate different project preferences, policies, politics, etc.
It seems like you might be more comfortable working on closed-source software with more like-minded engineers? You can have LLMs churn on all of your code 24/7, all without "weird politics" to deal with.
The trouble with asking questions is you sometimes get answers you don't wanna hear.
-- Nikki Sixx