besthvacaeo.agency Directory

Glossary

robots.txt for AI crawlers

A plain text file at the root of your site listing which crawlers may read it. The single most checkable thing in this entire subject.

Checkable · the test is below

What it means

Each AI company runs one or more named crawlers, and your robots.txt either permits or blocks each by name. It is a file. You can open it in a browser right now.

Why it matters for an HVAC company

Contractors have had these crawlers blocked by a previous developer, by a security plugin, or by advice to protect their content, and then paid for visibility work that could not possibly land. Checking costs nothing and it should be the first thing anyone does.

What would settle it

Visit yoursite.com/robots.txt and read it. Look for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended. Each is allowed or disallowed explicitly, and a blanket disallow covers them all.

A worked example

A contractor signs a 6-month retainer at $1,500 a month. Their robots.txt carries a blanket disallow added by a security plugin 2 years earlier. Four months and $6,000 of content work reached nothing, and one line in one file explained all of it. The check that would have caught it takes about 30 seconds.

The limit on that

Google-Extended governs training and Gemini grounding. It does not remove you from AI Overviews, which follow ordinary Google indexing.

ClaudeBot Retrieval augmented generation (RAG) Training data cutoff Seen in practice: the two OpenAI crawlers