TDD inside the agent loop - theater or actual value?
0 points by thang
0 points by thang
This posted a few hours before https://danluu.com/agentic-testing/ which also finds TDD doesn't help agents do better. But at least having tests before letting an agent go wild does make me feel better about the work I'm trying to hand over.
Based on Opus's judgment of the quality of the outcomes, ...
... Evaluation of adherence to TDD was also done by Sonnet 4.6.
Using LLMs to summaries and classify is playing to their strengths, but it feels circular to evaluate code with the same technique (albeit different model?). Garbage in, garbage out.