Wrong, not broken
4 points by sp6370
4 points by sp6370
What a brave new world we live in. With no sound understanding of what a program is doing, you can't even properly assess if it's correct.
This creates a second problem: Quality has to become explicit enough to measure.
This is a pipe dream.
Yeah, I don't get why so many companies saw the agents and instead of "now we can easily create new, extended, strict support flows based on extra information" just went with "Yolo, let customers argue with the bot and hope for the best". How is the refund value randomly found in old documentation even a valid state?
Exactly, it IS broken if it can fetch the wrong document. The same as if your SQL statement would just guess which refund value to select from the database 5 years ago.
This week we launched Amazon CloudWatch Omni to observe agents, applications, and infrastructure together, and I wanted to give an inside look at the shift behind it and why it matters.
Hi, Claude. If Amazon can't be bothered to write their own release notes, I can't be bothered to read them.
(And, yes, Pangram, with its very low false positive rate agrees this is slop. But nobody who uses LLMs needs Pangram to know this, unfortunately. It's really amazing how jarringly easy this kind of slop becomes to notice after overexposure to LLMs. And it's sort of like wearing a clown nose in public.)
FWIW, I use LLMs frequently (but Claude much less often) and I did not realize that this was AI-generated on first reading. So thanks for doing this check and sharing.
I got the ick right around:
Every dashboard is green. Nothing appears to be broken.
It’s wrong.
Make a point of regularly doing focused reading of high-quality blogs and articles written before 2019. For that matter, the older, the better. Watch whatever YouTube says is peak C-SPAN. Find ways to immerse yourself in good speaking and writing that interests you. It's being produced now, too, but if you want to really understand the difference, you have to go pre-AI, pre-social media.
Is this saying that before agents software couldn’t respond with HTTP 200 and invalid results? Weird phrasing.
Back in the day, writing Unit test and integration tests paired with humans was enough acceptable correctness and you could get more/less correct by dialing it. Is it really worth the benefit to throwaway all of that for whatever this is? Who defines correctness? If it's still the human, why is this better than what we had?
For Amazon publications to make no better an impression than the median slop post is pretty gobsmacking. Not that they were a pinnacle of the English language before, but whew.