Property Testing with Agent Swarms
4 points by ettolrach
4 points by ettolrach
I'm not doubting that an LLM can write property tests to find bugs, much like it can write unit tests to do the same.
But I always found that one of the core values of property based testing is that it can be a forcing function to think deeply about the core invariants of a system, and then capture those in an ideally smallish economical set of core properties that you can test. Letting LLMs do this work and come up with large numbers of properties seems to defeat the purpose at least a little bit, though I'm sure it can find bugs to fix.
The linked property testing skill has a lot of decent advice (compared to e.g. the trailofbits property testing skill linked to from How well do agents use test and verification techniques?). Their anecdata aside, I'm not claiming how well it might work for LLMs (it's very verbose, but lacking in specific examples), but it might be worth skimming for a human that's wondering how they might use property testing.
For context, I use property testing all the time; but I've seen others struggle to either come up with statements to test, or who write them in an ineffective way. I think a lot of it comes from the very basic examples used in tutorials, like x + y == y + x, get k (put k v db) == Just v, etc.. Real business functions rarely fulfill meaningful properties for completely arbitrary inputs (other than type-like properties, like "not empty", etc.). Where property testing shines is in generalising what would otherwise be unit tests: e.g. instead of setting up some particular situation, and asserting the exact result, we can step back and ask "what makes this situation interesting?" (e.g. "some of the rows have the same timestamp"), and "what aspect of the result is important?" (e.g. "no duplicates"). The bulletpoints in that skill are essentially brainstorming the sorts of systems/modules that might benefit from such scrutiny; the sorts of situations that might be interesting; and the sorts of invariants the results might be expected to satisfy.