I've Been Deindexed by Google
41 points by brn
41 points by brn
I was not expecting this to be a happy post just based on the title. Glad the author got what they wanted.
I am curious whether there is a drop in other server activity (LLM scrapers, other search engines, SSH login attempts, etc) after the Google index.
Same here. I would seriously doubt the scraping activity would drop, because I've got a few URLs/domains that have never been indexed by any search engine and are constantly hit by scrapers and bots attempting to find something (they look for .env files and .php files though there's no PHP running in that server).
I believe the list of publicly certified domain certificates from Let's Encrypt is the most popular source for bots now, since I've had new domains get hit within 1h of existing and the only "thing" that knew about them at that point was their service.
I believe the list of publicly certified domain certificates from Let's Encrypt is the most popular source for bots now, since I've had new domains get hit within 1h of existing and the only "thing" that knew about them at that point was their service.
Solution could be to do wildcard certs and then put stuff you don't want found on a subdomain.
Interestingly, 6% of my nobody-cares-website are coming from Kagi, whereas only <1% are from Google.
I accidentally blocked GoogleBot a few months ago and thought long and hard before white-listing it again. I get very few referrals from Google.
Long gone are the days where I would go out of my way to make my site easy to spider.
At my previous job we've had issues with Huawei bot trying to scrape absolutely everything from our server - we've done everything by the book (including their specifoc recommendations) and at the end of thr day we had to forcefully block its UA, as they claimed they update cached robots.txt every SIX MONTHS, so 3 months are really good!
I'm not sure I understand the end goal perfectly here.
The idea, basically, is that if you are not available via Google, you are harder to find by the scrapers? Because those are the annoying scrapper you don't want?
It’s really a shame that Google is so slow to update low-traffic websites. It should be way faster!
The only spiders I allow on my little web project are from the Internet Archive and from Marginalia Search.
I've been meaning to allow Kagi for a while but haven't gotten around to it.