Blog

Google-Extended Does Not Control AI Overviews

Manas Tripathi 9 min read
Share

Google-Extended governs use of content in Gemini, not in AI Overviews. AI Overviews and AI Mode are part of Search and crawled by Googlebot, so limiting what appears there means using snippet controls rather than blocking that agent. Providers now run separate training and answering crawlers, which means blocking the wrong one removes you from AI answers without affecting training. Cloudflare is reported to change its defaults for mixed-use crawlers on 15 September 2026, including for existing free-plan customers.

A great many Indian businesses added a line to their site’s crawler directives in the last two years, believing it opted them out of Google’s AI answers.

It did not. Google-Extended governs use of your content in Gemini and related products. AI Overviews and AI Mode are part of Search, and Search is crawled by Googlebot. Limiting what appears in those features is done with snippet directives, not by blocking that agent.

Which means a firm that added the block believing it had opted out has changed nothing about AI Overviews, and may believe the matter is settled.

That single misunderstanding is worth the whole article, but it is a symptom of a larger one.

Two jobs, two sets of agents

The instruction circulating through 2026 has been simple: block the AI bots.

It treats every AI crawler as the same thing, and they are not. The providers now run separate agents for separate jobs, and the distinction decides the outcome.

Training agents collect content to build models. OpenAI’s GPTBot is the best known. Google-Extended sits in the same category for Gemini. Anthropic and the others operate their own.

Answering agents fetch pages so a system can cite them in a response to a live question. OpenAI’s OAI-SearchBot serves ChatGPT’s search results. The other providers run equivalents.

Some are user-triggered — fetched because a person asked the assistant to look at a specific page.

Block the training agents and you have made a position on model training, which is a legitimate thing to have a position on.

Block the answering agents and you have removed yourself from the citations described in our article on who gets cited, while the training question is largely unaffected — because most training data was collected before you decided, and because the material is available from other copies of the web.

So the blunt instruction gets it precisely backwards. It surrenders the visibility and does not retrieve the content.

What we would actually do

A directive pointed at the wrong product

For an ordinary Indian business — a services firm, a manufacturer, an SME selling to other businesses — the position we would take is straightforward.

Allow the answering agents. You want to be cited. That is the entire objective of everything else in this cluster, and blocking the mechanism that delivers it makes no sense.

Decide about training agents on principle rather than on strategy. There is no visibility argument either way. If you object to your content being used to train commercial models without payment, block them and know that the practical effect on models already trained is close to nil. If you do not object, allowing them costs you nothing measurable.

Leave the user-triggered agents alone. Blocking those stops a person who explicitly asked their assistant to read your page. There is no version where that helps you.

The businesses with a genuinely different calculation are publishers whose content is the product, and firms with contractual obligations about where client material may travel. If your business is neither, the decision is less agonising than the discourse suggests.

The default changing on 15 September

Now the part that may take the decision out of your hands.

Cloudflare is reported to be changing its defaults on 15 September 2026, blocking mixed-use AI crawlers — agents that combine search, training and agent functions — by default on pages carrying advertising. The change is reported to apply to new customers, to new sites created by existing customers, and to all existing free-plan customers.

Read that last clause carefully if you are an Indian SME, because a very large number of sites here run on Cloudflare’s free plan. It was set up by a developer, it has not been touched since, and nobody in the business knows the settings exist.

The consequence is that a default somebody else chose could remove you from AI answers without any decision being made by you, and without anything visibly changing on your website.

We are not going to tell you whether Cloudflare’s policy is right. Publishers have a real grievance about content being used without compensation, and the same company’s payment programme has reportedly shifted toward paying when content is used in answers rather than when it is fetched, which is a more sensible basis.

What we will tell you is to go and look at your settings before that date rather than after, because finding out in November that you disappeared in September is an expensive way to learn what plan you are on.

The traffic split, which complicates our own advice

Two columns of agent names by job

One reported figure is worth putting against the recommendation above, because it makes the trade less comfortable than we have made it sound.

By June 2026, training crawlers reportedly made up around half of all AI bot traffic on Cloudflare’s network, while search crawlers were down to roughly a tenth.

Read that as a ratio. For every fetch made to answer somebody’s live question, several are made to collect training material. So a site that allows everything is mostly supplying training data and occasionally being cited.

That does not overturn the advice — the citations are the thing you want and there is no way to have them without allowing the agent that produces them — but it does mean the exchange is more one-sided than “allow the answering agents” implies. We would rather say that than present a clean recommendation.

It also explains why the infrastructure companies are moving. A ratio like that is a business problem for anyone whose content is expensive to produce, and it is why the compensation programmes have appeared.

If your content is the product

Everything above assumes an ordinary business that sells something other than words. Publishers, media companies and research firms have a different calculation and should not take the recommendation above.

For them the content is the asset, the citations frequently substitute for the visit rather than causing it, and the volume of training collection is a direct transfer of value with nothing returning.

The developments worth their attention are not in the text file. They are in the compensation programmes — notably the shift reported at Cloudflare from paying per fetch toward paying when content is actually used in an AI answer, which is a considerably better basis for a publisher than being paid for a crawl that may lead nowhere.

If you run a publication in India, the question to be working on is which of those marketplaces to enter and on what terms, not which line to add to a file. That is a commercial negotiation and it is above the level this article operates at.

Permission is only half of being fetchable

Allowing a crawler does nothing if there is nothing for it to collect.

The other half is rendering. Content injected into the page by JavaScript after load is frequently invisible to these systems, for the reason set out in our landing page article: they retrieve the page and do not run the scripts.

The check takes ten seconds. View the page source — the raw document, not the rendered inspector — and search for a sentence from the middle of your content. If it is not there, your permissions are irrelevant.

This catches a surprising number of Indian business sites, particularly those built on page builders and single-page frameworks by development shops optimising for how the site looks in a demo.

And it is the same failure as being absent from the index, which we cover in pages Google read and declined to keep — permission granted, content unavailable, nobody notices.

A directive is a request

A default setting changing on a dated calendar

Worth stating plainly, because a lot of business decisions are made on the assumption that these files are enforcement.

The crawler directive file is a convention. Well-behaved operators respect it. It is not a lock, it has no legal force in itself, and there have been public disputes through 2025 and 2026 about whether particular AI operators honour it.

If your objection to AI training is a matter of principle, the file expresses it and that has value.

If your objection is that you need to actually prevent access, the file is the wrong instrument, and the real answers are at the infrastructure level — which is precisely what the Cloudflare change is about, and why a company that sits in front of a large share of the web changing its defaults matters more than any line any of us adds to a text file.

One related caution. This file is also the single most dangerous file on your website. A misplaced directive can remove your site from search entirely, and it is the kind of error that ships during a rebuild and goes unnoticed for weeks — which is why it appears in the audit after a rebuild.

Do not edit it casually, do not let it be edited casually, and check it after any site change.

What we cannot tell you

We cannot tell you the full current list of agent names. Providers add and rename them, and any list published today is partly wrong within months. The taxonomy above is the durable part; the specific names need checking when you implement.

We cannot tell you what blocking training does to your position in models already built. Nobody outside those companies can, and claims in either direction are speculation.

And we cannot tell you what happens after 15 September, because the change had not taken effect when this was written. We will update this page after it does.

Final thoughts

Separate the two questions before you touch anything. Do you want to be cited, and do you object to being trained on. They are different decisions and they have different agents attached.

Then check three things this week: which directives your site currently carries, whether your content is in the raw page source, and which hosting plan you are on.

The third one has a date on it.

If you want us to check what your site is currently permitting and what changes on that date, you can reach out to us on whatsapp at +91 7738844851 .

More in Blog

Ready to talk about your growth?

Tell us what's stuck and we'll tell you what we'd do first. Free, 30 minutes, no pitch.

Want this done for your brand?
Work with us
Book
Link copied