Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

33% Positive

Analyzed from 489 words in the discussion.

Trending Topics

#training#cloudflare#google#pay#crawl#blocked#search#nothing#access#googlebot

Discussion (14 Comments)Read Original on HackerNews

simonwabout 2 hours ago
The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini:

> Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line with our call for transparency for website owners. Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service).

jofzarabout 1 hour ago
We had googlebot blast a random customer system and almost cause an outage, this is when I first learnt that google will use it for AI training also. It's honestly kind of frustrating also because you then search on it and theres (was) nothing on how you are meant to "correctly" tell google to fuck off, and not use it like that.
20k17 minutes ago
Google's web scraping functionality has been acting as a ddos for more than two decades. I've seen literally hundreds of reports of them attacking websites and taking them down, where there's nothing you can do but accept the traffic, or get delisted

This is unfortunately nothing new. There's no correct way to tell them to fuck off, they do not care, and they never will do. People have even taken them to court over this

inigyouabout 1 hour ago
Good. Cloudflare is a cancer on the internet. I want sites that use Cloudflare to disappear from Google, that'll teach their operators not to use it.
Cider998629 minutes ago
Why should I use something other than Cloudflare pages for a simple app landing page?
ajmurmann27 minutes ago
Why is this?
ceejayoz21 minutes ago
It's a planet-scale MITM?
tekacs44 minutes ago
> For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.

It's kind of exhausting seeing Cloudflare playing both sides of the arms race.

I just can't imagine bringing myself to use their technology to build agents and build AI products when they're also doing things like this.

tekacs41 minutes ago
> This also lines up the incentive model we want to foster. Losing trusted status across the more than 20% of web domains that sit behind Cloudflare is a deterrent with teeth. Trust becomes something you can carry with you, and something you can lose.

And even more so, LLM language aside, fun and fascinating to see them flagrantly calling out their position here as if it's a positive.

graemeabout 2 hours ago
Has there been any update on the pay per crawl program?
ray_vabout 2 hours ago
So, in summary: still the honors system. Got it. thanks.
zx8080about 1 hour ago
What's the "honors system"?
strictnein28 minutes ago
Robots.txt
zzzeek19 minutes ago
this is annoying, it makes a big deal about "Back when we announced pay-per-crawl"...

I want pay-per-crawl. I clicked the link for it a year ago, got presented with a "request access" button, I "requested access" and obviously since I'm nobody I heard absolutely nothing. Now they're touting the link again, I checked, still that same "request access" button. I have no idea if anyone even has access to this feature.

I don't care about all this other stuff, I want the AI crawlers to pay me cash. Because boy do those fuckers want to crawl me. I'll gladly double the size of my gerrit/jenkins servers to keep up with the load if these stupid bots want to pay to crawl every jenkins build artifact and every changeset source file on the server, as they really seem to want to do.