Howdy, Stranger!

It looks like you're new here. If you want to get involved, click one of these buttons!


BMail.ag - Secure Email Service
Server.net
CPLicense.net
VPS Server
Buy VPN
Vultr
VMs for AI
HostDare
ReliableSite White-Label Dedicated Hosting for Resellers
25% Recurring Discount on NVMe VPS
Try EnsoVPN - Reliable VPN - 1-Day Free Trial
InterServer VPS
BMail.ag - Secure Email Service
Best VPN
High-Performance Bare Metal Server Solutions
Karvl.com
Server Mania Cloud Hosting
DataWagon Hosting
AlphaVPS Hosting
Evoxt.com
Clouvider
VPS Hosting with NVMe
Residential IPs in the US & 4G Mobile Proxies in EU & US with Unlimited Bandwidth
ReliableSite White-Label Dedicated Hosting for Resellers
Rabisu - Hosting Solutions
CloudLinux
Try EnsoVPN - Fast & Private VPN - 1-Day Free Trial
New on LowEndTalk? Please Register and read our Community Rules.

All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.

Disappearing from Google Search Results

slowserversslowservers Member, Host Rep

This is about SporeStack, not Slow Servers. Slow Servers has always performed terribly in search engine results!

SporeStack does okay on Brave search and is at the top of DuckDuckGo for "monero vps."

However, it seems to be delisted on Google. I believe it used to be at the top for the same query, or very nearly.

Admittedly, this might be my fault because of my robots.txt, attempting to block all sorts of AI-related bots. Now I specifically didn't block the usual "GoogleBot" or whatever it's called. But maybe this is impacting my SEO in another way?

I'd appreciate any advice you can offer.

Thank you!

Comments

  • MrRadicMrRadic Host Rep, Veteran

    Why are you blocking AI bots? These deliver significantly higher quality traffic.

  • Pulsar67Pulsar67 Member, Host Rep

    I would probably check Google Search console and see if there are any issues.

  • rpqurpqu Member

    In doubt, reverse it

    Thanked by 1buggedout
  • trewqtrewq Administrator, Patron Provider

    @MrRadic said:
    Why are you blocking AI bots? These deliver significantly higher quality traffic.

    This. The amount of people using AI for research and decision making is off the charts, it’s almost the default.

    Google still drives traffic but it’s mostly people confirming their decision or just straight typing your business name to get to your website.

    Thanked by 2rpqu buggedout
  • rpqurpqu Member
    edited July 29

    @trewq said:

    @MrRadic said:
    Why are you blocking AI bots? These deliver significantly higher quality traffic.

    This. The amount of people using AI for research and decision making is off the charts, it’s almost the default.

    Google still drives traffic but it’s mostly people confirming their decision or just straight typing your business name to get to your website.

    Right, if the search proxy can't return result, the AI will try to curl it through my box. But, if it can't, it will give up, unless I intervened and manually copy paste it.
    Not all bots are scrapers. It could be your customer riding the clanker

    Thanked by 1buggedout
  • aphexaphex Member
    edited July 29

    @slowservers said: Admittedly, this might be my fault because of my robots.txt, attempting to block all sorts of AI-related bots. Now I specifically didn't block the usual "GoogleBot" or whatever it's called. But maybe this is impacting my SEO in another way?

    Google is multipurpose, if you block training you will also block results generally

    I've noticed it first sends feelers via GogoleBot then shortly after checks with Extended or others too and if rejected the whole thing is considered rejected

    @trewq said: This. The amount of people using AI for research and decision making is off the charts, it’s almost the default.

    This is a very opinionated provider that would rather explicitly not have AI as customers, which I respect

  • Yeah, the blocking of bots is going to be a strong challenge. Really the entire way that the internet generates revenue is broken and needs to change... and that is a whole other "loaded" topic haha.

    -you need to remove the sitewide Disallow: / (unless you truly want to hide the site from compliant crawlers too... which based on your query, you don't want. The broad block is a problem.

    -replace the blanket blocking with narrower rules for truly private paths only

    -leave googlebot unblocked, and make sure CSS, JS, and all page assets needed for rendering are crawlable.

    -use noindex for pages you do not want indexed rather than the robots.txt file if your goal is search exclusion.

    -re-submit the sitemap and request reindexing in search console after the fix.

    If you want to block AI scrapers, do it at the firewall/CDN layer if needed instead of a blanket robots.txt block

    Thanked by 1rpqu
  • zedzed Member

    that's the most fantastic robots.txt i've ever seen

    Thanked by 1slowservers
  • forestforest Member
    edited 12:25AM

    @trewq said: This. The amount of people using AI for research and decision making is off the charts, it’s almost the default.

    Unfortunate but true. These are not the same bots as the nasty scrapers that ruin the web for training, and they won't even obey robots.txt in the first place.

    Thanked by 1buggedout
  • rpqurpqu Member
    edited 12:29AM

    @zed said:
    that's the most fantastic robots.txt i've ever seen

    I can't archive it, so here's posted for archive
    User-agent: AddSearchBot
    User-agent: AI2Bot
    User-agent: AI2Bot-DeepResearchEval
    User-agent: Ai2Bot-Dolma
    User-agent: aiHitBot
    User-agent: amazon-kendra
    User-agent: Amazonbot
    User-agent: AmazonBuyForMe
    User-agent: Andibot
    User-agent: Anomura
    User-agent: anthropic-ai
    User-agent: Applebot
    User-agent: Applebot-Extended
    User-agent: atlassian-bot
    User-agent: Awario
    User-agent: bedrockbot
    User-agent: bigsur.ai
    User-agent: Bravebot
    User-agent: Brightbot 1.0
    User-agent: BuddyBot
    User-agent: Bytespider
    User-agent: CCBot
    User-agent: Channel3Bot
    User-agent: ChatGLM-Spider
    User-agent: ChatGPT Agent
    User-agent: ChatGPT-User
    User-agent: Claude-SearchBot
    User-agent: Claude-User
    User-agent: Claude-Web
    User-agent: ClaudeBot
    User-agent: Cloudflare-AutoRAG
    User-agent: CloudVertexBot
    User-agent: cohere-ai
    User-agent: cohere-training-data-crawler
    User-agent: Cotoyogi
    User-agent: Crawl4AI
    User-agent: Crawlspace
    User-agent: Datenbank Crawler
    User-agent: DeepSeekBot
    User-agent: Devin
    User-agent: Diffbot
    User-agent: DuckAssistBot
    User-agent: Echobot Bot
    User-agent: EchoboxBot
    User-agent: FacebookBot
    User-agent: facebookexternalhit
    User-agent: Factset_spyderbot
    User-agent: FirecrawlAgent
    User-agent: FriendlyCrawler
    User-agent: Gemini-Deep-Research
    User-agent: Google-CloudVertexBot
    User-agent: Google-Extended
    User-agent: Google-Firebase
    User-agent: Google-NotebookLM
    User-agent: GoogleAgent-Mariner
    User-agent: GoogleOther
    User-agent: GoogleOther-Image
    User-agent: GoogleOther-Video
    User-agent: GPTBot
    User-agent: iAskBot
    User-agent: iaskspider
    User-agent: iaskspider/2.0
    User-agent: IbouBot
    User-agent: ICC-Crawler
    User-agent: ImagesiftBot
    User-agent: imageSpider
    User-agent: img2dataset
    User-agent: ISSCyberRiskCrawler
    User-agent: Kangaroo Bot
    User-agent: KlaviyoAIBot
    User-agent: KunatoCrawler
    User-agent: laion-huggingface-processor
    User-agent: LAIONDownloader
    User-agent: LCC
    User-agent: LinerBot
    User-agent: Linguee Bot
    User-agent: LinkupBot
    User-agent: Manus-User
    User-agent: meta-externalagent
    User-agent: Meta-ExternalAgent
    User-agent: meta-externalfetcher
    User-agent: Meta-ExternalFetcher
    User-agent: meta-webindexer
    User-agent: MistralAI-User
    User-agent: MistralAI-User/1.0
    User-agent: MyCentralAIScraperBot
    User-agent: netEstate Imprint Crawler
    User-agent: NotebookLM
    User-agent: NovaAct
    User-agent: OAI-SearchBot
    User-agent: omgili
    User-agent: omgilibot
    User-agent: OpenAI
    User-agent: Operator
    User-agent: PanguBot
    User-agent: Panscient
    User-agent: panscient.com
    User-agent: Perplexity-User
    User-agent: PerplexityBot
    User-agent: PetalBot
    User-agent: PhindBot
    User-agent: Poggio-Citations
    User-agent: Poseidon Research Crawler
    User-agent: QualifiedBot
    User-agent: QuillBot
    User-agent: quillbot.com
    User-agent: SBIntuitionsBot
    User-agent: Scrapy
    User-agent: SemrushBot-OCOB
    User-agent: SemrushBot-SWA
    User-agent: ShapBot
    User-agent: Sidetrade indexer bot
    User-agent: Spider
    User-agent: TavilyBot
    User-agent: TerraCotta
    User-agent: Thinkbot
    User-agent: TikTokSpider
    User-agent: Timpibot
    User-agent: TwinAgent
    User-agent: VelenPublicWebCrawler
    User-agent: WARDBot
    User-agent: Webzio-Extended
    User-agent: webzio-extended
    User-agent: wpbot
    User-agent: WRTNBot
    User-agent: YaK
    User-agent: YandexAdditional
    User-agent: YandexAdditionalBot
    User-agent: YouBot
    User-agent: ZanistaBot
    Disallow: /
    Disallow: /secretscripts
    Allow: /scripts

  • forestforest Member

    @rpqu said:

    @zed said:
    that's the most fantastic robots.txt i've ever seen

    I can't archive it, so here's posted for archive

    You can do:

    ```
    line2
    line2
    line3
    ```
    

    To display it as:

    line1
    line2
    line3
    
  • rpqurpqu Member

    @forest said:

    @rpqu said:

    @zed said:
    that's the most fantastic robots.txt i've ever seen

    I can't archive it, so here's posted for archive

    You can do:

    > ```
    > line2
    > line2
    > line3
    > ```
    > 

    To display it as:

    line1
    line2
    line3
    

    I know, but the usual triple grave accent doesn't work.

  • forestforest Member
    edited 2:19AM

    @rpqu said:

    @forest said:

    @rpqu said:

    @zed said:
    that's the most fantastic robots.txt i've ever seen

    I can't archive it, so here's posted for archive

    You can do:

    > > ```
    > > line2
    > > line2
    > > line3
    > > ```
    > > 

    To display it as:

    line1
    line2
    line3
    

    I know, but the usual triple grave accent doesn't work.

    Huh, why isn't it working?

    (test)

    User-agent: AddSearchBot
    User-agent: AI2Bot
    User-agent: AI2Bot-DeepResearchEval
    User-agent: Ai2Bot-Dolma
    User-agent: aiHitBot
    User-agent: amazon-kendra
    User-agent: Amazonbot
    User-agent: AmazonBuyForMe
    User-agent: Andibot
    User-agent: Anomura
    User-agent: anthropic-ai
    User-agent: Applebot
    User-agent: Applebot-Extended
    User-agent: atlassian-bot
    User-agent: Awario
    User-agent: bedrockbot
    User-agent: bigsur.ai
    User-agent: Bravebot
    User-agent: Brightbot 1.0
    User-agent: BuddyBot
    User-agent: Bytespider
    User-agent: CCBot
    User-agent: Channel3Bot
    User-agent: ChatGLM-Spider
    User-agent: ChatGPT Agent
    User-agent: ChatGPT-User
    User-agent: Claude-SearchBot
    User-agent: Claude-User
    User-agent: Claude-Web
    User-agent: ClaudeBot
    User-agent: Cloudflare-AutoRAG
    User-agent: CloudVertexBot
    User-agent: cohere-ai
    User-agent: cohere-training-data-crawler
    User-agent: Cotoyogi
    User-agent: Crawl4AI
    User-agent: Crawlspace
    User-agent: Datenbank Crawler
    User-agent: DeepSeekBot
    User-agent: Devin
    User-agent: Diffbot
    User-agent: DuckAssistBot
    User-agent: Echobot Bot
    User-agent: EchoboxBot
    User-agent: FacebookBot
    User-agent: facebookexternalhit
    User-agent: Factset_spyderbot
    User-agent: FirecrawlAgent
    User-agent: FriendlyCrawler
    User-agent: Gemini-Deep-Research
    User-agent: Google-CloudVertexBot
    User-agent: Google-Extended
    User-agent: Google-Firebase
    User-agent: Google-NotebookLM
    User-agent: GoogleAgent-Mariner
    User-agent: GoogleOther
    User-agent: GoogleOther-Image
    User-agent: GoogleOther-Video
    User-agent: GPTBot
    User-agent: iAskBot
    User-agent: iaskspider
    User-agent: iaskspider/2.0
    User-agent: IbouBot
    User-agent: ICC-Crawler
    User-agent: ImagesiftBot
    User-agent: imageSpider
    User-agent: img2dataset
    User-agent: ISSCyberRiskCrawler
    User-agent: Kangaroo Bot
    User-agent: KlaviyoAIBot
    User-agent: KunatoCrawler
    User-agent: laion-huggingface-processor
    User-agent: LAIONDownloader
    User-agent: LCC
    User-agent: LinerBot
    User-agent: Linguee Bot
    User-agent: LinkupBot
    User-agent: Manus-User
    User-agent: meta-externalagent
    User-agent: Meta-ExternalAgent
    User-agent: meta-externalfetcher
    User-agent: Meta-ExternalFetcher
    User-agent: meta-webindexer
    User-agent: MistralAI-User
    User-agent: MistralAI-User/1.0
    User-agent: MyCentralAIScraperBot
    User-agent: netEstate Imprint Crawler
    User-agent: NotebookLM
    User-agent: NovaAct
    User-agent: OAI-SearchBot
    User-agent: omgili
    User-agent: omgilibot
    User-agent: OpenAI
    User-agent: Operator
    User-agent: PanguBot
    User-agent: Panscient
    User-agent: panscient.com
    User-agent: Perplexity-User
    User-agent: PerplexityBot
    User-agent: PetalBot
    User-agent: PhindBot
    User-agent: Poggio-Citations
    User-agent: Poseidon Research Crawler
    User-agent: QualifiedBot
    User-agent: QuillBot
    User-agent: quillbot.com
    User-agent: SBIntuitionsBot
    User-agent: Scrapy
    User-agent: SemrushBot-OCOB
    User-agent: SemrushBot-SWA
    User-agent: ShapBot
    User-agent: Sidetrade indexer bot
    User-agent: Spider
    User-agent: TavilyBot
    User-agent: TerraCotta
    User-agent: Thinkbot
    User-agent: TikTokSpider
    User-agent: Timpibot
    User-agent: TwinAgent
    User-agent: VelenPublicWebCrawler
    User-agent: WARDBot
    User-agent: Webzio-Extended
    User-agent: webzio-extended
    User-agent: wpbot
    User-agent: WRTNBot
    User-agent: YaK
    User-agent: YandexAdditional
    User-agent: YandexAdditionalBot
    User-agent: YouBot
    User-agent: ZanistaBot
    Disallow: /
    Disallow: /secretscripts
    Allow: /scripts
    

    Edit: It's working for me.

  • rpqurpqu Member

    @forest said:

    @rpqu said:

    @forest said:

    @rpqu said:

    @zed said:
    that's the most fantastic robots.txt i've ever seen

    I can't archive it, so here's posted for archive

    You can do:

    > > > ```
    > > > line2
    > > > line2
    > > > line3
    > > > ```
    > > > 

    To display it as:

    line1
    line2
    line3
    

    I know, but the usual triple grave accent doesn't work.

    Huh, why isn't it working?

    (test)

    User-agent: AddSearchBot
    User-agent: AI2Bot
    User-agent: AI2Bot-DeepResearchEval
    User-agent: Ai2Bot-Dolma
    User-agent: aiHitBot
    User-agent: amazon-kendra
    User-agent: Amazonbot
    User-agent: AmazonBuyForMe
    User-agent: Andibot
    User-agent: Anomura
    User-agent: anthropic-ai
    User-agent: Applebot
    User-agent: Applebot-Extended
    User-agent: atlassian-bot
    User-agent: Awario
    User-agent: bedrockbot
    User-agent: bigsur.ai
    User-agent: Bravebot
    User-agent: Brightbot 1.0
    User-agent: BuddyBot
    User-agent: Bytespider
    User-agent: CCBot
    User-agent: Channel3Bot
    User-agent: ChatGLM-Spider
    User-agent: ChatGPT Agent
    User-agent: ChatGPT-User
    User-agent: Claude-SearchBot
    User-agent: Claude-User
    User-agent: Claude-Web
    User-agent: ClaudeBot
    User-agent: Cloudflare-AutoRAG
    User-agent: CloudVertexBot
    User-agent: cohere-ai
    User-agent: cohere-training-data-crawler
    User-agent: Cotoyogi
    User-agent: Crawl4AI
    User-agent: Crawlspace
    User-agent: Datenbank Crawler
    User-agent: DeepSeekBot
    User-agent: Devin
    User-agent: Diffbot
    User-agent: DuckAssistBot
    User-agent: Echobot Bot
    User-agent: EchoboxBot
    User-agent: FacebookBot
    User-agent: facebookexternalhit
    User-agent: Factset_spyderbot
    User-agent: FirecrawlAgent
    User-agent: FriendlyCrawler
    User-agent: Gemini-Deep-Research
    User-agent: Google-CloudVertexBot
    User-agent: Google-Extended
    User-agent: Google-Firebase
    User-agent: Google-NotebookLM
    User-agent: GoogleAgent-Mariner
    User-agent: GoogleOther
    User-agent: GoogleOther-Image
    User-agent: GoogleOther-Video
    User-agent: GPTBot
    User-agent: iAskBot
    User-agent: iaskspider
    User-agent: iaskspider/2.0
    User-agent: IbouBot
    User-agent: ICC-Crawler
    User-agent: ImagesiftBot
    User-agent: imageSpider
    User-agent: img2dataset
    User-agent: ISSCyberRiskCrawler
    User-agent: Kangaroo Bot
    User-agent: KlaviyoAIBot
    User-agent: KunatoCrawler
    User-agent: laion-huggingface-processor
    User-agent: LAIONDownloader
    User-agent: LCC
    User-agent: LinerBot
    User-agent: Linguee Bot
    User-agent: LinkupBot
    User-agent: Manus-User
    User-agent: meta-externalagent
    User-agent: Meta-ExternalAgent
    User-agent: meta-externalfetcher
    User-agent: Meta-ExternalFetcher
    User-agent: meta-webindexer
    User-agent: MistralAI-User
    User-agent: MistralAI-User/1.0
    User-agent: MyCentralAIScraperBot
    User-agent: netEstate Imprint Crawler
    User-agent: NotebookLM
    User-agent: NovaAct
    User-agent: OAI-SearchBot
    User-agent: omgili
    User-agent: omgilibot
    User-agent: OpenAI
    User-agent: Operator
    User-agent: PanguBot
    User-agent: Panscient
    User-agent: panscient.com
    User-agent: Perplexity-User
    User-agent: PerplexityBot
    User-agent: PetalBot
    User-agent: PhindBot
    User-agent: Poggio-Citations
    User-agent: Poseidon Research Crawler
    User-agent: QualifiedBot
    User-agent: QuillBot
    User-agent: quillbot.com
    User-agent: SBIntuitionsBot
    User-agent: Scrapy
    User-agent: SemrushBot-OCOB
    User-agent: SemrushBot-SWA
    User-agent: ShapBot
    User-agent: Sidetrade indexer bot
    User-agent: Spider
    User-agent: TavilyBot
    User-agent: TerraCotta
    User-agent: Thinkbot
    User-agent: TikTokSpider
    User-agent: Timpibot
    User-agent: TwinAgent
    User-agent: VelenPublicWebCrawler
    User-agent: WARDBot
    User-agent: Webzio-Extended
    User-agent: webzio-extended
    User-agent: wpbot
    User-agent: WRTNBot
    User-agent: YaK
    User-agent: YandexAdditional
    User-agent: YandexAdditionalBot
    User-agent: YouBot
    User-agent: ZanistaBot
    Disallow: /
    Disallow: /secretscripts
    Allow: /scripts
    

    Edit: It's working for me.

    Must be encoding/clipboard problem

  • slowserversslowservers Member, Host Rep

    @aphex said:

    @slowservers said: Admittedly, this might be my fault because of my robots.txt, attempting to block all sorts of AI-related bots. Now I specifically didn't block the usual "GoogleBot" or whatever it's called. But maybe this is impacting my SEO in another way?

    Google is multipurpose, if you block training you will also block results generally

    I've noticed it first sends feelers via GogoleBot then shortly after checks with Extended or others too and if rejected the whole thing is considered rejected

    Oh man, that's what I was afraid of. SporeStack isn't shrinking, but I get the feeling that it may not be doing as well as it could be.

    The irony is that when I started SporeStack, I geared it towards ephemeral architecture and I specifically talked about autonomous systems being able to purchase servers. That was back in 2017! At the time, I think the thought was more around DAOs on Ethereum. Now, there's a very large group of capable AI agents that could actually be using it quite easily, which I'm probably blocking the majority of.

    I guess at the time, the idea seemed cool. I like it less and less as time goes on.

    My pickle is that I've setup SporeStack in a way that makes it very relevant to modern users, except that I have nerfed such usage to some extent. SporeStack, in some ways, makes sense to be as open market as possible.

    The irony is that one of my customers (a very nice one) is selling AI agents off SporeStack servers, despite my efforts to deter such activity. I've even been emailed by an OpenClaws agent!

    @TylerPYTHON said:
    Yeah, the blocking of bots is going to be a strong challenge. Really the entire way that the internet generates revenue is broken and needs to change... and that is a whole other "loaded" topic haha.

    -you need to remove the sitewide Disallow: / (unless you truly want to hide the site from compliant crawlers too... which based on your query, you don't want. The broad block is a problem.

    -replace the blanket blocking with narrower rules for truly private paths only

    -leave googlebot unblocked, and make sure CSS, JS, and all page assets needed for rendering are crawlable.

    -use noindex for pages you do not want indexed rather than the robots.txt file if your goal is search exclusion.

    -re-submit the sitemap and request reindexing in search console after the fix.

    If you want to block AI scrapers, do it at the firewall/CDN layer if needed instead of a blanket robots.txt block

    Actually, there's really nothing that shouldn't be crawled -- there's no private paths.

    I had a "scripts" and "secretscripts" folder. Both had AI poison code in them. I am not sure if anything actually went into "secretscripts" or not. I figured that the worst bots might.

    @Pulsar67 said:
    I would probably check Google Search console and see if there are any issues.

    I don't have that setup. I'm not sure if I want to bother or not, but maybe it'd be worth it. I think I had it going at one point. I hate dealing with the megacorps.

    --

    Another quandry is that I could allow money from AI sources to come in, while not distributing it back to AI causes. This is tricky for SporeStack, with Vultr and DigitalOcean both being very AI aggressive. It's also not something I want to make money from, but then again I think a lot of SporeStack might fall under that category.

    Even with Slow Servers, it's being used for VPNs specifically so people can have access to better AI models. My best efforts at making the least AI compatible infrastructure possible have not worked that well. So it seems that I'm both making some money off of it, but not as much as I could be. I am not a grand profiteer and a possible heretic at the end of the day.

    Will I go the way of the Dodo bird? Will I compromise on my values? Am I already?

    I do not know...

    It's harder to follow a hard line with a family, when there's more consequences for not going with the flow.

  • TylerPYTHONTylerPYTHON Member
    edited 4:19AM

    @slowservers said: Actually, there's really nothing that shouldn't be crawled -- there's no private paths.

    I had a "scripts" and "secretscripts" folder. Both had AI poison code in them. I am not sure if anything actually went into "secretscripts" or not. I figured that the worst bots might.

    I get the anti-AI angle, honestly. But the thing is, blocking bots with robots.txt mostly just hurts your own site traffic and visibility. The bad actor bots usually don’t care about robots.txt anyway and will just ignore it.

    The bigger issue is that Disallow: / is blocking the good bots too, so Google and other legit crawlers can’t really see or understand the public site properly. That’s probably doing more damage than good.

    If the goal is to keep the AI scrapers out without tanking search traffic, I’d look at a managed bot detection service at the CDN layer instead. That gives you a much better shot at blocking the junk while still letting the legit search bots through.

    Thanked by 2forest rpqu
  • forestforest Member

    @slowservers said: I've even been emailed by an OpenClaws agent!

    Lmao, what was the email about?

Sign In or Register to comment.